CourseLarge Language Models · Module 12: Multimodal Models · part 64 of 80
Part 64 · Module 12: Multimodal Models

Part 3: Other modalities

25 min read·22 Sept 2026

Voice messages: audio input and transcription

Some Brightlane customers, especially on mobile, record a voice message instead of typing. There are two ways to handle audio:

  • Speech-to-text (ASR, automatic speech recognition) turns audio into a transcript, which then flows into the text pipeline you already have (triage, retrieval, drafting). Whisper is the best-known ASR model family; Groq serves whisper-large-v3 and whisper-large-v3-turbo through the same OpenAI-compatible API as its chat models.
  • Speech-native models accept audio directly in the chat request (next section).

Before choosing, know what audio files weigh. Uncompressed PCM audio (what a WAV file holds) is sample rate x channels x bytes per sample per second. Speech APIs recommend 16 kHz mono; Groq's speech-to-text page suggests converting to 16 kHz mono before upload, and Google's audio page says Gemini downsamples audio to 16 Kbps and merges channels to mono anyway.

python
"""Voice messages on tickets: sizes, durations, speech segments, cost, and transcription.

Everything except transcribe() and ask_about_audio() runs offline. transcribe() calls
Groq's OpenAI-compatible transcription endpoint (GROQ_API_KEY); ask_about_audio() sends
the audio to a Gemini chat model (GEMINI_API_KEY). Both run only with M12_LIVE=1.
Note: voice_note.wav holds tones, not speech, so a real transcript of it is
meaningless; record your own voice note to try the live call.
Run: PYTHONPATH=. python examples/m12_audio.py
"""
from __future__ import annotations

import base64
import math
import os
import struct
import wave
from pathlib import Path

from supportdesk.llm import Usage, chat, make_client
from supportdesk.pricing import cost_usd

WAV = Path("data/attachments/voice_note.wav")
# console.groq.com/docs/speech-to-text, checked 21 Sep 2026: USD per hour, 10 s minimum per request
WHISPER_PRICES = {"whisper-large-v3": 0.111, "whisper-large-v3-turbo": 0.04}
MIN_BILLED_S = 10
GEMINI_AUDIO_TOKENS_PER_S = {"audio-understanding page": 32, "media-resolution page, Gemini 3": 25}


def wav_info(path: Path) -> dict:
    with wave.open(str(path), "rb") as w:
        frames, rate = w.getnframes(), w.getframerate()
        return {"seconds": frames / rate, "rate": rate, "channels": w.getnchannels(),
                "bits": 8 * w.getsampwidth(), "bytes": path.stat().st_size}


def speech_segments(path: Path, frame_ms: int = 20, threshold: float = 0.02, min_gap_ms: int = 200) -> list[tuple[float, float]]:
    """Energy-based voice activity detection: the simplest endpointing a voice agent can do."""
    with wave.open(str(path), "rb") as w:
        rate, raw = w.getframerate(), w.readframes(w.getnframes())
    samples = struct.unpack(f"<{len(raw) // 2}h", raw)
    step = rate * frame_ms // 1000
    active = []
    for i in range(0, len(samples) - step + 1, step):
        chunk = samples[i:i + step]
        rms = math.sqrt(sum(s * s for s in chunk) / len(chunk)) / 32768
        active.append(rms > threshold)
    segments, start, silent = [], None, 0
    for k, on in enumerate(active + [False] * (min_gap_ms // frame_ms + 1)):
        if on:
            start = k if start is None else start
            silent = 0
        elif start is not None:
            silent += 1
            if silent * frame_ms >= min_gap_ms:
                segments.append((start * frame_ms / 1000, (k - silent + 1) * frame_ms / 1000))
                start, silent = None, 0
    return segments


def whisper_cost(seconds: float, model: str) -> float:
    return max(seconds, MIN_BILLED_S) / 3600 * WHISPER_PRICES[model]


def chunk_plan(total_s: float, chunk_s: float = 600, overlap_s: float = 5) -> list[tuple[float, float]]:
    """Split long audio into overlapping chunks so no word is cut in half at a boundary."""
    plan, start = [], 0.0
    while start < total_s:
        plan.append((start, min(start + chunk_s, total_s)))
        start += chunk_s - overlap_s
    return plan


def transcribe(path: Path, model: str = "whisper-large-v3") -> dict:
    """Groq Whisper through the OpenAI client (same base URL and key as supportdesk.llm)."""
    client = make_client("groq")
    with path.open("rb") as f:
        result = client.audio.transcriptions.create(
            model=model, file=f, response_format="verbose_json", temperature=0.0)
    return {"text": result.text, "language": getattr(result, "language", None),
            "duration": getattr(result, "duration", None)}


def ask_about_audio(path: Path, question: str, model: str = "gemini-3.5-flash") -> str:
    """Speech-native path: the audio goes straight into a chat model (Gemini's OpenAI-compatible input_audio)."""
    audio_b64 = base64.b64encode(path.read_bytes()).decode("ascii")
    messages = [{"role": "user", "content": [
        {"type": "text", "text": question},
        {"type": "input_audio", "input_audio": {"data": audio_b64, "format": "wav"}},
    ]}]
    return chat(messages, provider="gemini", model=model, max_tokens=300).text


def main() -> None:
    info = wav_info(WAV)
    print(f"{WAV.name}: {info['seconds']:.1f} s, {info['rate']} Hz, {info['channels']} ch, "
          f"{info['bits']}-bit, {info['bytes']:,} bytes")
    for label, rate, ch, bits in [("16 kHz mono 16-bit (speech APIs)", 16_000, 1, 16),
                                  ("44.1 kHz stereo 16-bit (phone recorder)", 44_100, 2, 16)]:
        print(f"  one minute as {label}: {rate * ch * bits // 8 * 60 / 1e6:.2f} MB")

    segs = speech_segments(WAV)
    print(f"\nSpeech segments found by energy VAD: {len(segs)}")
    print("  " + ", ".join(f"{a:.2f}-{b:.2f}s" for a, b in segs))

    print("\nTranscription cost per voice note (Groq Whisper, 10 s minimum):")
    for seconds in (info["seconds"], 45, 180):
        row = ", ".join(f"{m} ${whisper_cost(seconds, m):.6f}" for m in WHISPER_PRICES)
        print(f"  {seconds:6.1f} s: {row}")
    monthly = 3000 * whisper_cost(45, "whisper-large-v3")
    print(f"  3,000 notes/month at 45 s each with whisper-large-v3: ${monthly:.2f}")

    print("\nSending the same audio to a speech-native LLM instead (input side only):")
    for source, rate in GEMINI_AUDIO_TOKENS_PER_S.items():
        tokens = math.ceil(45 * rate)
        usd = cost_usd(Usage(input_tokens=tokens), "gemini-3.5-flash")
        print(f"  45 s at {rate} tokens/s ({source}): {tokens:,} tokens, ${usd:.6f} at pricing.py's input rate")

    plan = chunk_plan(45 * 60)
    print(f"\nA 45-minute call recording in 10-minute chunks with 5 s overlap: {len(plan)} chunks")
    print("  " + ", ".join(f"{a / 60:.1f}-{b / 60:.1f} min" for a, b in plan))

    if os.environ.get("M12_LIVE") == "1":
        print("\nLive transcription:", transcribe(WAV))
        print("Speech-native answer:", ask_about_audio(
            WAV, "Is this a person speaking? Describe the recording and the speaker's tone in two sentences."))


if __name__ == "__main__":
    main()

Code explained

  • In simple words: measure the voice note, find where the sound is, price transcribing it two ways, plan how to split long recordings, and (with keys) transcribe it or ask a model about it.
  • What happens:
    • wav_info() reads the header with the standard wave module: frames, rate, channels, sample width.
    • speech_segments() is energy-based voice activity detection (VAD): split the audio into 20 ms frames, compute each frame's loudness (root mean square), mark frames above a threshold as active, and merge active runs separated by less than 200 ms of quiet. This is the simplest possible endpointing.
    • whisper_cost() applies Groq's per-hour prices (0.111 USD for whisper-large-v3, 0.04 USD for turbo, checked 21 September 2026) with the documented 10-second minimum per request.
    • chunk_plan() splits long audio into 10-minute chunks with 5 seconds of overlap so no word is cut at a boundary. Groq's free tier caps uploads at 25 MB; a 45-minute 16 kHz mono WAV is about 86 MB.
    • transcribe() calls client.audio.transcriptions.create() on Groq's client from make_client(), asking for verbose_json, which adds segment timestamps and the detected language.
    • ask_about_audio() sends the audio as an input_audio content part to a Gemini chat model, the format Google's OpenAI-compatibility page documents.
    • main() prints everything and runs the live calls only with M12_LIVE=1.
  • Comes out:

With a Groq key, the live transcription is one call:

bash
export GROQ_API_KEY=your-key-here GEMINI_API_KEY=your-other-key
M12_LIVE=1 python examples/m12_audio.py

Code explained

  • In simple words: the same script, plus a real transcription and a real speech-native answer.
  • What happens: transcribe() uploads the WAV to Groq's /audio/transcriptions endpoint; ask_about_audio() sends it to Gemini as input_audio.
  • Comes out: Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.

    text
    Live transcription: {'text': ' Thank you.', 'language': 'English', 'duration': 5.2}
    Speech-native answer: This is not a person speaking. It is a series of evenly spaced electronic tones, about a second apart, with no words or vocal tone.
    

    Our voice note is tones, not speech, so the only correct transcript is empty. The illustrative "Thank you." shows a real, documented failure mode: ASR models can hallucinate words over non-speech audio. In "Careless Whisper" (Koenecke et al., FAccT 2024, arXiv 2402.08021), roughly 1% of Whisper transcriptions contained entire phrases or sentences that were not in the audio, and hallucinations were more common for speakers with long non-vocal pauses. Two defenses follow. Run VAD first and do not transcribe audio without speech (our energy VAD cannot tell tones from speech, so production systems use a trained VAD model such as Silero VAD). And keep the audio attached to the ticket so an agent can check a transcript that matters. To test transcription for real, record yourself reading ticket T-1001 aloud on your phone, save it as WAV or M4A, and point WAV at it.

Speech-native models

A speech-native model takes audio tokens directly, the same way a VLM takes image tokens, instead of reading a transcript. Audio is cut into short frames (Gemini documents 25 to 32 tokens per second), encoded, and placed in the context. Because the model hears the audio rather than a transcript, it can use tone, hesitation, background noise, and non-speech sounds, and it can answer questions about the recording ("does the caller sound angry?", "is there music in the background?"). Real-time voice models go one step further and also generate audio tokens, so one model listens and speaks (examples current in September 2026: OpenAI's gpt-realtime-2.1 over WebRTC or WebSocket, and Google's Live API over WebSockets).

SituationUse thisWhy
Voice notes on tickets that feed the existing text pipelineASR (Whisper) then textCheapest per minute, the transcript is searchable, loggable, and redactable, and every later step is unchanged
You need tone, emotion, or non-speech soundsA speech-native modelA transcript throws that information away
Long calls you must search or audit laterASR with timestamps (verbose_json)Timestamps let an agent jump to the moment in question
Live conversation with a callerA real-time speech-to-speech model, or a streaming ASR plus LLM plus TTS pipelineSee the latency budget below
Regulated data where you must show exactly what was saidASR plus keeping the audioA transcript is auditable text; a speech-native model's reasoning is not

Video: frame sampling

Video understanding is mostly image understanding repeated. Models do not watch video continuously: they sample frames at some rate (Gemini's default is 1 frame per second), turn each frame into image tokens, and usually add the audio track as audio tokens. The cost is frames x tokens per frame + seconds x audio tokens per second, which makes the sampling rate the biggest lever you have. Customers send screen recordings of bugs, and a screen recording is mostly the same frame over and over.

python
"""Video as a stack of images: frame sampling math and real frame extraction.

The math section is exact arithmetic on published per-frame token counts.
The extraction section samples the synthetic 60 s screen recording (a GIF, so
Pillow can read it without ffmpeg) and drops near-duplicate frames.
Run: PYTHONPATH=. python examples/m12_video_frames.py
"""
from __future__ import annotations

import json
import math
from pathlib import Path

from PIL import Image, ImageChops, ImageSequence, ImageStat

from examples.m12_image_tokens import claude_tokens, openai_patch_tokens
from supportdesk.llm import Usage
from supportdesk.pricing import cost_usd

GIF = Path("data/attachments/screen_recording.gif")
# Per-frame tokens from Google's docs (checked 21 Sep 2026): video-understanding page (66 at low,
# the default there; 258 at high; audio 32 tokens/s) and media-resolution page for Gemini 3
# (70 at default/low/medium; 280 at high; audio 25 tokens/s). Values: (tokens per frame, audio tokens/s).
GEMINI_PER_FRAME = {"video page low": (66, 32), "video page high": (258, 32),
                    "Gemini 3 default": (70, 25), "Gemini 3 high": (280, 25)}


def frames_at(path: Path, fps: float) -> list[tuple[float, Image.Image]]:
    """Frames at t = 0, 1/fps, 2/fps, ... using each GIF frame's duration to find what is on screen.

    Times are kept in integer milliseconds: summing 0.2 s durations as floats drifts, and a frame
    that starts at 47.0 s would be found at 46.99999 s, one frame too early.
    """
    timeline, t_ms = [], 0
    with Image.open(path) as gif:
        for frame in ImageSequence.Iterator(gif):
            dur_ms = int(frame.info.get("duration", 100))
            timeline.append((t_ms, t_ms + dur_ms, frame.convert("L").copy()))
            t_ms += dur_ms
    count = math.floor(t_ms * fps / 1000)
    picks = [round(k * 1000 / fps) for k in range(count)]
    return [(ms / 1000, next(img for a, b, img in timeline if a <= ms < b)) for ms in picks]


def drop_near_duplicates(frames: list[tuple[float, Image.Image]], threshold: float = 1.0) -> list[tuple[float, Image.Image]]:
    """Keep a frame only if its mean absolute pixel difference from the last kept frame exceeds threshold."""
    kept = [frames[0]]
    for ts, img in frames[1:]:
        diff = ImageStat.Stat(ImageChops.difference(img, kept[-1][1])).mean[0]
        if diff > threshold:
            kept.append((ts, img))
    return kept


def main() -> None:
    seconds = 60
    print(f"Token math for a {seconds} s clip with its audio track (frames x per-frame + seconds x audio rate):")
    print(f"{'fps':>5s} {'frames':>7s} " + " ".join(f"{k:>18s}" for k in GEMINI_PER_FRAME))
    for fps in (0.25, 0.5, 1, 2, 5):
        n = int(seconds * fps)
        cells = [f"{n * per + seconds * audio:>18,d}" for per, audio in GEMINI_PER_FRAME.values()]
        print(f"{fps:>5} {n:>7d} " + " ".join(cells))
    n1 = seconds  # 1 fps
    print(f"\nSame 60 frames sent as separate 1280x720 images instead: OpenAI patch {n1 * openai_patch_tokens(1280, 720):,}"
          f" tokens, Claude {n1 * claude_tokens(1280, 720):,} tokens")
    per, audio = GEMINI_PER_FRAME["Gemini 3 high"]
    usd = cost_usd(Usage(input_tokens=n1 * per + seconds * audio), "gemini-3.5-flash")
    print(f"1 fps, Gemini 3 high, on gemini-3.5-flash at pricing.py's input rate: ${usd:.4f} per minute of video")

    truth = json.loads(Path("data/attachments/ground_truth.json").read_text())["screen_recording.gif"]
    print(f"\nReal extraction from {GIF.name} ({truth['seconds']} s, screens: {truth['screens']})")
    for fps in (0.2, 1, 5):
        sampled = frames_at(GIF, fps)
        kept = drop_near_duplicates(sampled)
        print(f"  sample at {fps} fps: {len(sampled):3d} frames, after dropping near-duplicates: {len(kept):2d} "
              f"at t = {', '.join(f'{ts:g}' for ts, _ in kept)}")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: compute what a one-minute video costs at several frame rates, then actually sample our screen recording and throw away frames that look the same as the last one we kept.
  • What happens:
    • GEMINI_PER_FRAME holds Google's published per-frame and per-second audio numbers from two pages that disagree slightly (the video page: 66 tokens per frame at low resolution, its default, or 258 at high, with 32 audio tokens per second; the Gemini 3 media-resolution page: 70 per frame at default, 280 at high, 25 audio tokens per second).
    • frames_at() walks the GIF's frames using each frame's duration to build a timeline, then picks whatever frame is on screen at t = 0, 1/fps, 2/fps, ....
    • drop_near_duplicates() compares each sampled frame with the last kept frame by mean absolute pixel difference (Pillow's ImageChops.difference and ImageStat) and keeps it only if the difference exceeds a threshold of 1.0 (on a 0 to 255 scale).
    • main() prints the token table, the same 60 frames priced as separate images, the dollar cost of one minute, and the real extraction at three rates.
  • Comes out:

    text
    Token math for a 60 s clip with its audio track (frames x per-frame + seconds x audio rate):
      fps  frames     video page low    video page high   Gemini 3 default      Gemini 3 high
     0.25      15              2,910              5,790              2,550              5,700
      0.5      30              3,900              9,660              3,600              9,900
        1      60              5,880             17,400              5,700             18,300
        2     120              9,840             32,880              9,900             35,100
        5     300             21,720             79,320             22,500             85,500
    
    Same 60 frames sent as separate 1280x720 images instead: OpenAI patch 66,240 tokens, Claude 71,760 tokens
    1 fps, Gemini 3 high, on gemini-3.5-flash at pricing.py's input rate: $0.0274 per minute of video
    
    Real extraction from screen_recording.gif (60 s, screens: ['Board: Q3 Launch', 'Export > CSV clicked', 'Exporting... 38%', 'Error E-4012: export failed'])
      sample at 0.2 fps:  12 frames, after dropping near-duplicates:  8 at t = 0, 15, 25, 30, 35, 40, 45, 50
      sample at 1 fps:  60 frames, after dropping near-duplicates: 11 at t = 0, 12, 25, 28, 31, 34, 37, 40, 43, 46, 47
      sample at 5 fps: 300 frames, after dropping near-duplicates: 13 at t = 0, 12, 25, 27.4, 29.8, 32, 34.4, 36.8, 39.2, 41.4, 43.8, 46.2, 47
    

    One minute of video at Gemini's default 1 fps is 5,700 to 5,880 tokens at default resolution, and three times that at high resolution. Sending the same 60 frames as ordinary 1280x720 images to a size-based provider would cost 66,000 to 72,000 tokens, because native video paths use far fewer tokens per frame. On our recording, the screen changes at 0, 12, 25, and 47 seconds. At 1 fps, near-duplicate dropping keeps 11 of 60 frames: the four screen changes, each on time, plus frames of the progress bar growing between 25 and 47 seconds, which is "different" pixel-wise but not in meaning (a higher threshold or a crop that ignores the progress bar would drop those too). At 0.2 fps (one frame every 5 seconds) it sees each change late: the export click at 15 s instead of 12, and the error at 50 s instead of 47. Anything shorter than 5 seconds could be missed entirely. The lessons: sample at least as fast as the shortest event you care about, and deduplicate before you pay. For this recording, "1 fps, then drop near-duplicates" sends 11 frames instead of 60, about 3,080 image tokens at Gemini 3 high resolution instead of 16,800.

A failure worth knowing about: the first version of frames_at() summed the GIF's frame durations (mostly 0.2 s) as floating-point seconds. After more than 200 additions the error screen's start time was 47.000000000000185 instead of 47.0, so the sample at exactly 47 s fell inside the previous (exporting) frame, and the error screen first appeared at 48 s. The symptom was a timestamp one second late; the diagnosis was printing the timeline boundaries around 47 s; the fix was integer milliseconds. Timestamps are evidence in a support ticket ("the error appeared at 0:47"), so this kind of bug matters.

SituationUse thisWhy
Screen recording of a bug (mostly static)1 fps plus near-duplicate dropping, or scene-change detectionMeasured: 11 of 60 frames kept, every screen change caught
Fast action (a UI animation glitch, a flicker)Higher fps on a short clipped segment1 fps misses sub-second events
Long recordings where speech carries the content (a recorded call with screen share)Transcribe the audio, sample frames sparselyMost information is in the audio
You only need the final stateLast frame (or last few) as a single imageOne image costs 70 to 2,048 tokens instead of thousands

Image generation and editing, at a working level

Image generation models produce an image from a text prompt; image editing models change an existing image given a prompt, often with reference images or a mask (the region to change). As of September 2026 the main hosted families are Google's Gemini image models (gemini-3.1-flash-image, gemini-3-pro-image, gemini-3.1-flash-lite-image; gemini-2.5-flash-image is listed as legacy) and OpenAI's GPT Image models (gpt-image-2.5-sunburst and gpt-image-2.5-flare in the Images API and as a Responses API tool). Google's pricing page lists gemini-3.1-flash-image at about 0.067 USD per 1K-resolution image. Google states that all its image models embed a SynthID watermark.

For a support desk, generation is a minor use (illustrations for help-center articles). Editing is more interesting, but the most common edit Brightlane needs, redaction of personal data from screenshots before they are stored, logged, or sent to a third-party model (Module 11), should not be done by a generative model at all. A generative edit re-draws pixels and can alter or invent content. Redaction must be exact and repeatable, so we do it with OCR and a black rectangle.

python
"""Image generation (hosted model) and a deterministic edit (redaction) for Brightlane.

generate_illustration() needs GEMINI_API_KEY and M12_LIVE=1; it uses the image
generation example from Google's OpenAI-compatibility docs. redact() runs
offline: it finds text with OCR and paints over it, exactly and repeatably.
Run: PYTHONPATH=. python examples/m12_image_edit.py
"""
from __future__ import annotations

import base64
import os
import re
from pathlib import Path

from PIL import Image, ImageDraw
from rapidocr_onnxruntime import RapidOCR

from supportdesk.llm import make_client

ATT = Path("data/attachments")
# The model in Google's OpenAI-compatibility example (checked 21 Sep 2026). Google lists it as legacy;
# newer image models (gemini-3.1-flash-image, gemini-3-pro-image) are documented on the native API.
IMAGE_MODEL = os.environ.get("M12_IMAGE_MODEL", "gemini-2.5-flash-image")


def generate_illustration(prompt: str, out: Path) -> Path:
    """Text-to-image through Gemini's OpenAI-compatible Images endpoint."""
    client = make_client("gemini")
    response = client.images.generate(model=IMAGE_MODEL, prompt=prompt, response_format="b64_json", n=1)
    out.write_bytes(base64.b64decode(response.data[0].b64_json))
    return out


def redact(path: Path, patterns: list[str], out: Path) -> list[str]:
    """Black out every OCR line that matches a pattern. Returns the texts removed."""
    engine = RapidOCR(intra_op_num_threads=1, inter_op_num_threads=1)
    img = Image.open(path).convert("RGB")
    result, _ = engine(img)
    draw, removed = ImageDraw.Draw(img), []
    for box, text, _ in result or []:
        if any(re.search(p, text) for p in patterns):
            xs, ys = [p[0] for p in box], [p[1] for p in box]
            draw.rectangle([min(xs) - 3, min(ys) - 3, max(xs) + 3, max(ys) + 3], fill="black")
            removed.append(text)
    img.save(out)
    return removed


def main() -> None:
    out = ATT / "error_dialog_redacted.png"
    removed = redact(ATT / "error_dialog.png", [r"Request\s*ID"], out)
    print(f"redacted {len(removed)} line(s): {removed}")
    check = redact(out, [r"[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}"], ATT / "error_dialog_recheck.png")
    print(f"re-scan of the redacted image finds request ids: {check or 'none'}")

    prompt = ("Flat illustration for a help-center article: a kanban board with three columns, "
              "one card highlighted, Brightlane blue (#406ebe) accents, no text, no logos.")
    if os.environ.get("M12_LIVE") == "1":
        print("saved", generate_illustration(prompt, ATT / "kb_illustration.png"))
    else:
        print(f"would call images.generate(model={IMAGE_MODEL!r}) with a {len(prompt)}-character prompt")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: black out any text line that matches a pattern, check the result by reading it again, and (with a key) generate a help-center illustration.
  • What happens:
    • redact() runs OCR, and for every recognized line matching one of the regex patterns, draws a filled rectangle a few pixels larger than the line's box. It returns what it removed, for the audit log.
    • main() redacts the request ID line, then runs a second pass on the redacted image looking for anything shaped like a request ID. That re-scan is the test: redaction you do not verify is redaction you hope for.
    • generate_illustration() calls client.images.generate() on Gemini's OpenAI-compatible endpoint with the model from Google's compatibility example, and decodes the base64 result. Override the model with M12_IMAGE_MODEL to try a newer one; availability through the compatibility layer varies by model, so check the error message if a call is refused.
  • Comes out: (about 4 seconds)

    text
    redacted 1 line(s): ['RequestID:7f3c-91ab-22de']
    re-scan of the redacted image finds request ids: none
    would call images.generate(model='gemini-2.5-flash-image') with a 155-character prompt
    

    OCR read the line without spaces (RequestID:7f3c-91ab-22de) and the regex still matched it because \s* allows zero spaces. The re-scan finds nothing. One limit to know: OCR-based redaction only removes text that OCR detects. Faint, tiny, or handwritten personal data can slip through, so for high-risk data combine it with blocking whole known regions (for example, the account menu in every Brightlane screenshot).

.
SituationUse thisWhy
Remove personal data from screenshotsOCR plus pattern matching plus rectangles, then a verification re-scanExact, repeatable, auditable
Illustrations for help articlesA hosted image generation modelFast, cheap, good enough for decoration; review before publishing
Change a product screenshot for docs (new button label)Re-capture the real UIA generated "screenshot" can show UI that does not exist
Mock-ups and conceptsImage editing with reference imagesIterative edits are where these models shine

Real-time voice: latency, interruption, and turn-taking

A voice agent on Brightlane's phone line has to decide when the caller has finished (endpointing or end-of-turn detection), transcribe, think, and start speaking, all fast enough to feel like a conversation. Human conversation sets a demanding bar: Levinson and Torreira (2015, Frontiers in Psychology) summarize the evidence that gaps between turns are "of the order of 200 ms", while producing speech takes people over 600 ms, so humans predict the end of a turn and start planning early. Stivers et al. (2009, PNAS) found the same pattern of avoiding overlap and minimizing silence across 10 languages, with average gaps differing by up to about 250 ms between languages.

A cascaded voice agent (ASR, then LLM, then text-to-speech) adds up its stages. Here is a budget built from published figures, with the two assumptions labelled, and a turn manager that handles barge-in (the caller starts talking while the agent is speaking).

python
"""Real-time voice: a latency budget with sourced figures, and barge-in handling.

The budget is arithmetic on published numbers (sources in STAGES) plus two
labelled assumptions. The turn manager is deterministic code driven by a
scripted event timeline, so its output is real.
Run: PYTHONPATH=. python examples/m12_voice_latency.py
"""
from __future__ import annotations

from dataclasses import dataclass, field

# (stage, milliseconds, source). Checked 21 Sep 2026; vendor figures exclude network overhead.
STAGES = [
    ("end-of-turn wait (silence endpointing)", 500, "OpenAI realtime VAD guide example: silence_duration_ms 500"),
    ("ASR transcript after end of turn", 150, "ElevenLabs docs: scribe_v2_realtime partials in ~150 ms"),
    ("LLM first chunk", 770, "Artificial Analysis: gpt-oss-120b on Groq, median first chunk 0.77 s"),
    ("LLM first sentence (15 tokens)", round(15 / 474 * 1000), "same page: 474 tokens/s output"),
    ("TTS time to first audio", 75, "ElevenLabs docs: eleven_flash_v2_5 '~75ms'"),
    ("network, 3 hops", 150, "assumption: about 50 ms per hop; measure yours"),
]
# Same page: with reasoning_effort high, the first *answer* token arrives after 4.99 s (reasoning first).
REASONING_FIRST_ANSWER_MS = 4990
HUMAN_GAP_MS = 200  # Levinson and Torreira 2015: gaps between turns are "of the order of 200 ms"


def budget(stages: list[tuple[str, int, str]], eager_saving_ms: int = 0) -> int:
    return sum(ms for _, ms, _ in stages) - eager_saving_ms


@dataclass
class TurnManager:
    """LISTENING -> THINKING -> SPEAKING, with barge-in: user speech during SPEAKING stops playback."""
    state: str = "LISTENING"
    history: list[dict] = field(default_factory=list)
    log: list[str] = field(default_factory=list)
    _reply: str = ""
    _played_chars: int = 0

    def on(self, t_ms: int, event: str, payload: str = "") -> None:
        if event == "user_final" and self.state in ("LISTENING", "THINKING"):
            self.history.append({"role": "user", "content": payload})
            self.state = "THINKING"
        elif event == "reply_ready" and self.state == "THINKING":
            self._reply, self._played_chars, self.state = payload, 0, "SPEAKING"
        elif event == "played" and self.state == "SPEAKING":
            self._played_chars = min(len(self._reply), self._played_chars + int(payload))
            if self._played_chars == len(self._reply):
                self._finish(interrupted=False)
        elif event == "user_speech" and self.state == "SPEAKING":
            self._finish(interrupted=True)  # barge-in: stop TTS now, keep only what was heard
        self.log.append(f"{t_ms:>6d} ms  {event:12s} -> {self.state}")

    def _finish(self, interrupted: bool) -> None:
        heard = self._reply[: self._played_chars]
        if interrupted:
            heard = heard.rsplit(" ", 1)[0] + " [interrupted]"
        self.history.append({"role": "assistant", "content": heard})
        self.state = "LISTENING"


def main() -> None:
    print(f"{'stage':40s} {'ms':>5s}  source")
    for name, ms, source in STAGES:
        print(f"{name:40s} {ms:>5d}  {source}")
    total = budget(STAGES)
    print(f"{'time to first audio (sequential)':40s} {total:>5d}")
    eager = budget(STAGES, eager_saving_ms=200)
    print(f"{'with eager end-of-turn (-200)':40s} {eager:>5d}  Deepgram: fires 150-250 ms earlier, "
          "50-70% more LLM calls")
    slow = budget(STAGES) - 770 + REASONING_FIRST_ANSWER_MS
    print(f"{'same, reasoning_effort high':40s} {slow:>5d}  first answer token 4.99 s on the same page")
    print(f"human gap between turns for comparison: about {HUMAN_GAP_MS} ms")

    print("\nBarge-in on a scripted call:")
    tm = TurnManager()
    reply = "Your export failed because the board has more than ten thousand cards. You can filter the board first."
    for t, event, payload in [(0, "user_final", "Why did my CSV export fail?"),
                              (1640, "reply_ready", reply),
                              (2200, "played", "30"), (2800, "played", "30"),
                              (3100, "user_speech", ""),
                              (3900, "user_final", "Can I export to JSON instead?")]:
        tm.on(t, event, payload)
    print("\n".join(tm.log))
    print("history the model sees next turn:")
    for m in tm.history:
        print(f"  {m['role']:9s} {m['content']}")


if __name__ == "__main__":
    main()

Code explained

  • In simple words: add up how long each step of a voice reply takes to see how far we are from a human pause, then run a tiny state machine that stops talking the moment the caller interrupts.
  • What happens:
    • STAGES lists each stage with its milliseconds and source: a 500 ms silence wait (the value in OpenAI's realtime VAD guide example), about 150 ms for a streaming ASR transcript (ElevenLabs documents partial transcripts in about 150 ms for scribe_v2_realtime), 0.77 s median first chunk and 474 tokens per second for gpt-oss-120b on Groq (Artificial Analysis provider page, checked 21 September 2026), about 75 ms to first audio for ElevenLabs eleven_flash_v2_5, and an assumed 50 ms per network hop. Vendor figures exclude network and application time.
    • budget() adds the stages, optionally subtracting an eager end-of-turn saving. Deepgram reports that its Flux model's EagerEndOfTurn fires 150 to 250 ms earlier than its normal end-of-turn signal, at the cost of 50 to 70% more LLM calls (you start generating speculatively and discard the reply if the caller keeps talking).
    • REASONING_FIRST_ANSWER_MS is the same Artificial Analysis page's time to the first answer token for gpt-oss-120b at high reasoning effort: 4.99 s, because a reasoning model thinks before it answers (Module 5).
    • TurnManager has three states. user_final moves to THINKING, reply_ready to SPEAKING, played advances how many characters of the reply have been spoken, and user_speech during SPEAKING is a barge-in: stop playback, and write into the history only the part the caller actually heard, cut at a word boundary and marked [interrupted].
    • main() prints the budget and replays a scripted call with an interruption. The timeline is scripted; the state machine's behavior is real.
  • Comes out:

    text
    stage                                       ms  source
    end-of-turn wait (silence endpointing)     500  OpenAI realtime VAD guide example: silence_duration_ms 500
    ASR transcript after end of turn           150  ElevenLabs docs: scribe_v2_realtime partials in ~150 ms
    LLM first chunk                            770  Artificial Analysis: gpt-oss-120b on Groq, median first chunk 0.77 s
    LLM first sentence (15 tokens)              32  same page: 474 tokens/s output
    TTS time to first audio                     75  ElevenLabs docs: eleven_flash_v2_5 '~75ms'
    network, 3 hops                            150  assumption: about 50 ms per hop; measure yours
    time to first audio (sequential)          1677
    with eager end-of-turn (-200)             1477  Deepgram: fires 150-250 ms earlier, 50-70% more LLM calls
    same, reasoning_effort high               5897  first answer token 4.99 s on the same page
    human gap between turns for comparison: about 200 ms
    
    Barge-in on a scripted call:
         0 ms  user_final   -> THINKING
      1640 ms  reply_ready  -> SPEAKING
      2200 ms  played       -> SPEAKING
      2800 ms  played       -> SPEAKING
      3100 ms  user_speech  -> LISTENING
      3900 ms  user_final   -> THINKING
    history the model sees next turn:
      user      Why did my CSV export fail?
      assistant Your export failed because the board has more than ten [interrupted]
      user      Can I export to JSON instead?
    

    About 1.7 seconds from the caller going quiet to the first audio, against a human gap of about 200 ms. The biggest items are the LLM's first chunk (770 ms) and the silence wait (500 ms), so those are where to work: a faster or smaller model, a shorter or smarter end-of-turn decision, streaming the first sentence to TTS as soon as it exists (already assumed here: only 15 tokens are waited for), and filler audio ("Let me check that") when a tool call is needed. Turning on high reasoning effort makes it 5.9 seconds, which is unusable for voice: use low reasoning effort or a non-reasoning model on the voice path. The barge-in history matters for correctness: the model's next turn sees that the caller heard only "...more than ten", not the advice to filter the board, so it should not assume the caller knows about filtering.

.

The three hard problems in turn-taking are all trade-offs between speed and false triggers:

SituationUse thisWhy
Deciding the caller has finishedA semantic or model-based end-of-turn detector (OpenAI's semantic_vad, Deepgram Flux) rather than silence aloneA 500 ms silence rule cuts off people who pause mid-sentence and waits too long for people who do not
Shaving latency furtherEager end-of-turn with speculative generation150 to 250 ms earlier, paid for with 50 to 70% more LLM calls (Deepgram's figures)
The caller interruptsStop TTS immediately, truncate the assistant's turn to what was heard, go back to listeningOtherwise the model believes the caller heard things they did not
Background noise or "uh-huh" while the agent talksRequire a minimum speech duration or a semantic check before treating sound as a barge-inBackchannels are not interruptions
Lowest latency and natural prosodyA speech-to-speech real-time modelNo separate ASR and TTS hops; less control and harder to log and test
Control, auditability, and choosing each componentA cascaded ASR plus LLM plus TTS pipelineEvery stage is swappable and every transcript is loggable