CourseLarge Language Models · Module 12: Multimodal Models · part 65 of 80
Part 65 · Module 12: Multimodal Models

Part 4: Multimodal system design

22 min read·22 Sept 2026

Choosing between a multimodal model and a specialized pipeline

Every attachment type has two families of solution: send it to a general multimodal model, or run a specialized pipeline (parser, OCR, ASR, frame sampler) and send text onward. The measurements in this module point to a consistent rule: use the specialized pipeline for the exact, high-volume, checkable parts, and the multimodal model for interpretation and for whatever the pipeline cannot handle.

SituationUse thisWhy
Exact identifiers in screenshots (error codes, request IDs)OCR plus regexMeasured 6/6 correct on exact strings, local, free, no hallucinated characters
"What went wrong in this screenshot?"VLMNeeds interpretation; no specialized tool does it
Born-digital PDFsText layer plus parser108/108 fields, 2 ms per document
Scans of known document typesOCR plus tolerant parser plus checks, VLM on failure91.7% of fields held-out, every wrong document flagged
Documents of unknown layoutVLM with schema plus checksOne prompt covers many layouts
Voice notesASR, then the text pipelineAbout 4 USD a month at 3,000 notes; transcripts are searchable and redactable
Tone, emotion, or non-speech soundSpeech-native modelTranscripts discard it
Screen recordingsSample, deduplicate, then VLM on the kept frames11 frames instead of 60 on our recording
Counting, measuring, exact positionsA specialized detector, or ask the customerVLM counting and spatial reasoning are documented weak spots
Data that must not leave your infrastructureLocal pipeline, or a local VLM through OllamaPrivacy beats convenience

Cost and latency of multimodal inputs

Here are this module's numbers side by side. Local times are from the course machine (2 CPU cores) and will differ on yours; token counts are exact under each provider's published rule; dollar figures use prices checked 21 September 2026.

InputReaderTokensCost per itemLocal time
Born-digital invoice PDFpypdf plus parser00about 2 ms
Scanned invoice (one page)RapidOCR plus parser00 (CPU only)about 1 s
Scanned invoice (one page)Gemini 3.5 Flash, default resolution1,120 image + about 150 textabout 0.0037 USD (with 200 output tokens)network plus model time
Screenshot 1280x800Claude Haiku 4.51,334 (504 at 768 px)0.0013 USD (0.0005)network plus model time
Screenshot, any sizeGroq Qwen 3.8 27B2,0480.0016 USDnetwork plus model time
45 s voice noteGroq whisper-large-v3not token-billed0.0014 USDnetwork plus model time
45 s voice noteGemini, speech-native1,125 to 1,440about 0.002 USD at pricing.py's input ratenetwork plus model time
60 s screen recording, 1 fpsGemini 3, default resolution5,700 including audioabout 0.009 USDnetwork plus model time
60 s screen recording, deduplicated11 frames at Gemini 3 highabout 3,080about 0.005 USDabout 90 ms to sample

Two patterns stand out. Per item, multimodal input is cheap; the costs that bite are volume (every frame, every agent step, every page of a long PDF) and latency (a model call is hundreds of milliseconds to seconds, while a parser takes milliseconds). And the context window fills fast: 60 frames at high resolution is 16,800 image tokens before a word of the prompt, so budgets from Module 2 and Module 7 need image-aware counting, which is what estimate_input_tokens() is for.

Evaluating multimodal outputs

Everything in Module 10 applies; multimodal inputs add a few specifics:

  • Generate or label ground truth per field. Synthetic assets (as here) give free, exact labels for mechanics; a few hundred labelled real attachments give the numbers you actually ship on.
  • Score fields, not documents only. Per-field accuracy shows which field breaks (v3's total and plan); per-document accuracy shows the business impact ("how many invoices need a person?"). Report both, with n and a confidence interval.
  • Hold out a test set and do not look at it while building. Measured here: 100% on dev, 91.7% held-out, 12 of 12 perfect documents versus 4 of 12.
  • Normalize before comparing. 12,480 and 12480 are the same answer; 288 and 288.00 are the same amount. Decide the normalization once, in code, and use it for every reader.
  • Count abstentions separately from errors. A reader that says "unreadable" or None is safer than one that guesses: its failures are routable. v4's held-out misses were all abstentions.
  • Measure what the checks catch. "Wrong documents flagged" and "false alarms" tell you whether auto-accepting checked outputs is safe.
  • Slice by input condition: layout, scan quality, image size, language, audio quality. The average hides the slice that fails.
  • For descriptions and summaries, where there is no exact answer, use Module 10's rubric-based LLM-as-judge, and give the judge the image too, or it grades fluency rather than faithfulness.
  • Re-run on every model or prompt change. Vision behavior changes between model versions more than text behavior does, and providers change resize rules without changing model names.

Tests for the building blocks

The tests pin down every deterministic piece: token formulas against the providers' examples, the resize caps, both parsers on every generated PDF, the checks, the VLM plumbing with ScriptedLLM, image-aware token estimation, the audio helpers, frame sampling timestamps, barge-in, and type sniffing.

tests/test_m12_multimodal.py

python
"""Tests for Module 12. Run: PYTHONPATH=. pytest -q tests/test_m12_multimodal.py
OCR-dependent tests are marked slow-ish (a few seconds each on CPU)."""
from __future__ import annotations

import json
import random
import subprocess
import sys
from pathlib import Path

import pytest

from examples.m12_audio import chunk_plan, speech_segments, whisper_cost
from examples.m12_image_tokens import (claude_resize, claude_tokens, gemini_legacy_tokens,
                                       openai_patch_tokens, openai_tile_tokens)
from examples.m12_invoice_pipeline import (FIELDS, KNOWN_ACCOUNTS, best_account, check_invoice, image_message,
                                           money, parse_v1, parse_v2, parse_v4, parse_vlm_json, pdf_text, snap,
                                           vlm_extract)
from examples.m12_voice_latency import STAGES, budget
from examples.m12_video_frames import drop_near_duplicates, frames_at
from examples.m12_vision_chat import estimate_input_tokens, screenshot_messages, encode_image
from examples.m12_voice_latency import TurnManager
from examples.m12_lab import sniff
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages

ATT = Path("data/attachments")


@pytest.fixture(scope="session", autouse=True)
def assets():
    if not (ATT / "ground_truth.json").exists():
        subprocess.run([sys.executable, "examples/m12_make_assets.py"], check=True)
    return json.loads((ATT / "ground_truth.json").read_text())


@pytest.mark.parametrize("got,want", [
    (openai_patch_tokens(1024, 1024), 1229), (openai_patch_tokens(2048, 2048), 3000),
    (openai_patch_tokens(4096, 512), 2458), (openai_tile_tokens(1024, 1024), 765),
    (openai_tile_tokens(2048, 4096), 1105), (openai_tile_tokens(4000, 3000, detail="low"), 85),
    (claude_tokens(200, 200), 64), (claude_tokens(1000, 1000), 1296), (claude_tokens(1092, 1092), 1521),
    (claude_tokens(1920, 1080), 1560), (claude_tokens(2000, 1500), 1564), (claude_tokens(3840, 2160), 1560),
    (claude_tokens(1920, 1080, "high"), 2691), (claude_tokens(3840, 2160, "high"), 4784),
    (gemini_legacy_tokens(960, 540), 1548), (gemini_legacy_tokens(384, 384), 258),
])
def test_token_formulas_match_provider_examples(got, want):
    assert got == want


def test_claude_resize_never_enlarges_and_respects_caps():
    rng = random.Random(0)
    for _ in range(300):
        w, h = rng.randint(10, 6000), rng.randint(10, 6000)
        rw, rh = claude_resize(w, h)
        assert rw <= w and rh <= h and max(rw, rh) <= 1568 and claude_tokens(rw, rh) <= 1568


def test_layout_parser_reads_every_generated_pdf(assets):
    for truth in assets["invoices"]:
        got = parse_v2(pdf_text(ATT / "invoices" / f"{truth['invoice_number']}.pdf", layout=True))
        assert {f: str(got[f]) for f in FIELDS} == {f: str(truth[f]) for f in FIELDS}


def test_naive_parser_fails_on_modern_layout(assets):
    modern = next(t for t in assets["invoices"] if t["layout"] == "modern")
    got = parse_v1(pdf_text(ATT / "invoices" / f"{modern['invoice_number']}.pdf"))
    assert all(v is None for v in got.values())


def test_checks_flag_arithmetic_errors(assets):
    good = dict(assets["invoices"][0])
    assert check_invoice(good) == []
    assert "subtotal + tax != total" in check_invoice({**good, "total": "289.00"})
    assert "missing plan" in check_invoice({**good, "plan": None})


def test_snap_repairs_ocr_spacing_but_not_strangers():
    assert snap("NorthwindStudio", KNOWN_ACCOUNTS) == "Northwind Studio"
    assert snap("AcmeRobotics.", KNOWN_ACCOUNTS) == "Acme Robotics"
    assert snap("Totally Unknown LLC", KNOWN_ACCOUNTS) == "Totally Unknown LLC"


def test_vlm_path_plumbing_with_stand_in():
    llm = ScriptedLLM(replies=['```json\n{"invoice_number": "INV-2026-004512", "seats": "24", "total": 288}\n```'])
    got = vlm_extract(ATT / "invoices" / "INV-2026-004512_scan.png", llm=llm)
    parts = llm.calls[0]["messages"][0]["content"]
    assert [p["type"] for p in parts] == ["text", "image_url"]
    assert parts[1]["image_url"]["url"].startswith("data:image/png;base64,")
    assert got["seats"] == 24 and got["total"] == "288.00" and got["customer"] is None
    assert parse_vlm_json("no json here") == {f: None for f in FIELDS}


def test_image_messages_are_counted_as_images_not_text():
    url, _, size = encode_image(ATT / "error_dialog.png", 1024)
    msgs = screenshot_messages(url)
    naive = count_messages(msgs)
    for provider in ("groq", "gemini", "ollama"):
        assert estimate_input_tokens(msgs, provider, [size]) < naive / 10
    assert image_message("hi", ATT / "usage_chart.png")[0]["content"][1]["type"] == "image_url"


def test_audio_helpers():
    assert len(speech_segments(ATT / "voice_note.wav")) == 6
    assert whisper_cost(3, "whisper-large-v3") == whisper_cost(10, "whisper-large-v3")
    plan = chunk_plan(2700)
    assert plan[0][0] == 0 and plan[-1][1] == 2700
    assert all(a[1] > b[0] for a, b in zip(plan, plan[1:]))  # consecutive chunks overlap


def test_frame_sampling_and_dedupe(assets):
    frames = frames_at(ATT / "screen_recording.gif", fps=1)
    assert len(frames) == assets["screen_recording.gif"]["seconds"]
    kept = drop_near_duplicates(frames)
    assert len(kept) < len(frames)
    kept_times = [ts for ts, _ in kept]
    assert {12.0, 25.0, 47.0} <= set(kept_times)  # every screen change is kept, on time


def test_barge_in_keeps_only_what_was_heard():
    tm = TurnManager()
    tm.on(0, "user_final", "hi")
    tm.on(10, "reply_ready", "one two three four five six")
    tm.on(20, "played", "9")
    tm.on(30, "user_speech")
    assert tm.state == "LISTENING"
    assert tm.history[-1]["content"] == "one two [interrupted]"


def test_sniff_uses_bytes_not_names(tmp_path):
    fake = tmp_path / "invoice.pdf"
    fake.write_bytes((ATT / "error_dialog.png").read_bytes())
    assert sniff(fake) == "image"
    assert sniff(ATT / "voice_note.wav") == "audio"
    assert sniff(ATT / "screen_recording.gif") == "video"


def test_money_takes_last_separator_as_decimal_point():
    assert money("Subtotal 552,00") == "552.00"
    assert money("Total due:1,320.00 USD") == "1320.00"
    assert money("Amiountdue    144:00") == "144.00"
    assert money("no amount here") is None


def test_v4_on_known_ocr_noise():
    text = ("INVOICE    No.    1NV-2026-004549\nBrightanenc    panssi:    2026-03-18\nCustomer    QuantaLabs'\n"
            "Teamplanx10users@12:00    120.00\nSubtotal    120.00\n.VAT/Tax.    24.00.\nAmiountdue    144:00\nCurrency    EUR:")
    got = parse_v4(text)
    assert got == {"invoice_number": "INV-2026-004549", "invoice_date": "2026-03-18", "customer": "Quanta Labs",
                   "plan": "Team", "seats": 10, "subtotal": "120.00", "tax": "24.00", "total": "144.00",
                   "currency": "EUR"}
    assert check_invoice(got) == []
    assert best_account("Totally Unknown LLC", KNOWN_ACCOUNTS) is None


def test_heldout_split_is_separate(assets):
    dev = {t["invoice_number"] for t in assets["invoices"]}
    held = {t["invoice_number"] for t in assets["invoices_heldout"]}
    assert len(dev) == len(held) == 12 and not dev & held
    for t in assets["invoices_heldout"]:
        assert (ATT / "invoices" / "heldout" / f"{t['invoice_number']}_scan.png").exists()


def test_latency_budget_adds_up():
    assert budget(STAGES) == sum(ms for _, ms, _ in STAGES)
    assert budget(STAGES, eager_saving_ms=200) == budget(STAGES) - 200

Code explained

  • In simple words: 31 fast checks that the arithmetic, parsers, and plumbing still behave after any change.
  • What happens: the assets session fixture regenerates the attachments if ground_truth.json is missing. test_token_formulas_match_provider_examples compares 16 token computations with the providers' published values; test_claude_resize_never_enlarges_and_respects_caps tries 300 random sizes; the parser tests read every generated PDF; test_vlm_path_plumbing_with_stand_in uses ScriptedLLM to check the request shape and type coercion (plumbing, not model quality); the rest cover the checks, money(), parse_v4(), the dev/held-out split, image-aware token estimates, audio helpers, frame timestamps, barge-in, the latency budget, and type sniffing.
  • Comes out: nothing by itself; run it with the command below.
bash
python -m pytest -q tests/test_m12_multimodal.py

Code explained

  • In simple words: run the suite.
  • What happens: the session fixture regenerates the attachments if ground_truth.json is missing. The parametrized test compares 16 token computations with the providers' published values. Other tests check that the resize never enlarges and respects caps on 300 random sizes, that v2 reads all 12 dev PDFs, that v1 fails on the modern layout, that the checks flag arithmetic errors, that money() and parse_v4() handle known OCR noise, that dev and held-out do not overlap, that image messages are counted as images, that the frame sampler keeps every screen change at the right second, that barge-in keeps only heard words, and that sniff() trusts bytes rather than file names. OCR-heavy checks use the cache, so the suite is fast.
  • Comes out:

Module Lab

Build the attachment router for ticket T-1001 ("Charged twice this month"). Maya's team wants every attachment on a ticket read by the cheapest reader that can be checked, with anything uncertain routed to a person, and the cost of any model escalation shown. The lab combines the type sniffing, PDF parsing, OCR, invoice checks, audio analysis, frame sampling, and cost estimates from this module.

examples/m12_lab.py

python
"""Module 12 lab: route every attachment on a Brightlane ticket to the cheapest reliable reader.

Real offline: type sniffing, PDF parsing, OCR, arithmetic checks, audio and video
analysis, token and cost estimates. With M12_LIVE=1 and a key, escalations go to a
vision model and voice notes to Whisper; otherwise they go to the human review queue.
Run: PYTHONPATH=. python examples/m12_lab.py
"""
from __future__ import annotations

import json
import os
import re
import time
from pathlib import Path

import pypdfium2 as pdfium
from PIL import Image

from examples.m12_audio import speech_segments, transcribe, wav_info, whisper_cost
from examples.m12_extraction_eval import cached_ocr
from examples.m12_image_tokens import gemini3_tokens
from examples.m12_invoice_pipeline import check_invoice, parse_v2, parse_v4, pdf_text, vlm_extract
from examples.m12_video_frames import drop_near_duplicates, frames_at
from examples.m12_vision_chat import VISION_MODELS
from supportdesk.data import load_tickets
from supportdesk.llm import Usage, resolve
from supportdesk.pricing import cost_usd

ATT = Path("data/attachments")
LIVE = os.environ.get("M12_LIVE") == "1"
VLM_PRICE_MODEL = "gemini-3.5-flash"  # used only to estimate what an escalation would cost


def sniff(path: Path) -> str:
    """Decide the type from the first bytes, never from the file name a customer chose."""
    head = path.read_bytes()[:12]
    if head.startswith(b"%PDF"):
        return "pdf"
    if head.startswith(b"\x89PNG") or head.startswith(b"\xff\xd8\xff"):
        return "image"
    if head.startswith(b"GIF8"):
        return "video"  # we treat animated GIFs as screen recordings
    if head[:4] == b"RIFF" and head[8:12] == b"WAVE":
        return "audio"
    return "unknown"


def vlm_estimate_usd(n_images: int = 1, output_tokens: int = 200) -> float:
    return cost_usd(Usage(input_tokens=n_images * gemini3_tokens("high") + 150, output_tokens=output_tokens),
                    VLM_PRICE_MODEL)


def handle_invoice_fields(fields: dict, source: Path, how: str) -> dict:
    issues = check_invoice(fields)
    if not issues:
        return {"route": how, "fields": fields, "status": "auto-accepted"}
    if LIVE:
        provider, _ = resolve()
        image = source if source.suffix == ".png" else rasterize(source)
        vlm_fields = vlm_extract(image, model=VISION_MODELS[provider])
        if not check_invoice(vlm_fields):
            return {"route": how + " -> VLM", "fields": vlm_fields, "status": "auto-accepted after escalation"}
    return {"route": how, "fields": fields, "status": f"human review ({'; '.join(issues)})",
            "escalation_estimate_usd": round(vlm_estimate_usd(), 5)}


def rasterize(pdf: Path, dpi: int = 100) -> Path:
    out = pdf.with_name(pdf.stem + "_page1.png")
    pdfium.PdfDocument(str(pdf))[0].render(scale=dpi / 72).to_pil().save(out)
    return out


def process(path: Path) -> dict:
    kind = sniff(path)
    started = time.perf_counter()
    if kind == "pdf":
        text = pdf_text(path, layout=True)
        if len(text.strip()) > 50:  # a real text layer: parse it, no model needed
            report = handle_invoice_fields(parse_v2(text), path, "pdf text layer")
        else:                       # image-only PDF: rasterize, OCR, tolerant parser
            report = handle_invoice_fields(parse_v4(cached_ocr(rasterize(path))), path, "scanned pdf: OCR")
    elif kind == "image" and "INVOICE" in cached_ocr(path).upper():  # a photographed or scanned invoice
        report = handle_invoice_fields(parse_v4(cached_ocr(path)), path, "image: OCR")
    elif kind == "image":
        text = cached_ocr(path)  # cheap local pass for exact identifiers
        codes = sorted(set(re.findall(r"\bE-\d{4}\b", text)))
        ids = re.findall(r"Request\s*ID:?\s*([0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4})", text)
        report = {"route": "OCR for ids" + (" + VLM for description" if LIVE else ""),
                  "fields": {"error_codes": codes, "request_ids": ids},
                  "status": "ids extracted" if codes else "human review (no error code found)",
                  "vlm_description_estimate_usd": round(vlm_estimate_usd(), 5)}
    elif kind == "audio":
        info = wav_info(path)
        report = {"route": "speech-to-text", "fields": {"seconds": round(info["seconds"], 1),
                                                        "speech_segments": len(speech_segments(path))},
                  "status": "transcribed" if LIVE else "queued for transcription",
                  "transcription_usd": round(whisper_cost(info["seconds"], "whisper-large-v3"), 6)}
        if LIVE:
            report["fields"]["transcript"] = transcribe(path)["text"]
    elif kind == "video":
        kept = drop_near_duplicates(frames_at(path, fps=1))
        report = {"route": "1 fps sampling + near-duplicate drop",
                  "fields": {"frames_kept": len(kept), "at_seconds": [ts for ts, _ in kept]},
                  "status": "frames ready for VLM",
                  "vlm_estimate_usd": round(vlm_estimate_usd(n_images=len(kept)), 5)}
    else:
        report = {"route": "rejected", "fields": {}, "status": "unsupported type"}
    report.update(file=str(path.relative_to(ATT)), kind=kind,
                  local_ms=round((time.perf_counter() - started) * 1000))
    return report


def main() -> None:
    ticket = next(t for t in load_tickets() if t.id == "T-1001")
    attachments = [ATT / "invoices/INV-2026-004512.pdf", ATT / "invoices/INV-2026-004514_scanned.pdf",
                   ATT / "invoices/heldout/INV-2026-005528_scan.png", ATT / "error_dialog.png",
                   ATT / "voice_note.wav", ATT / "screen_recording.gif"]
    print(f"{ticket.id}: {ticket.subject!r} with {len(attachments)} attachments (live={LIVE})\n")
    reports = [process(p) for p in attachments]
    for r in reports:
        print(f"{r['file']}  [{r['kind']}] via {r['route']}  ({r['local_ms']} ms local)")
        print(f"    status: {r['status']}")
        print(f"    fields: {json.dumps(r['fields'])}")
        extra = {k: v for k, v in r.items() if k.endswith("_usd")}
        if extra:
            print(f"    cost:   {extra}")
    auto = sum(r["status"].startswith(("auto", "ids", "frames", "transcribed")) for r in reports)
    spend = sum(v for r in reports for k, v in r.items() if k.endswith("_usd"))
    print(f"\nhandled without a person: {auto}/{len(reports)}; model spend if every estimate were used: ${spend:.4f}")
    Path("data/attachments/lab_report.json").write_text(json.dumps(reports, indent=2))


if __name__ == "__main__":
    main()

Code explained

  • In simple words: a mail room for attachments: look at what each file really is, send it to the right desk, check the desk's work, and only pay a model (or a person) when the checks fail.
  • What happens:
    • sniff() decides the type from the file's first bytes (magic numbers: %PDF, the PNG signature, GIF8, RIFF....WAVE), never from the file name or extension a customer chose. Module 11's rule: attachments are untrusted input.
    • handle_invoice_fields() runs check_invoice(). Clean results are auto-accepted. Otherwise, in live mode it escalates to a VLM and accepts that result only if it passes the same checks; offline, it routes to human review with the reasons and the estimated cost of escalating.
    • rasterize() renders an image-only PDF page with pypdfium2 so OCR (or a VLM) can read it.
    • process() routes each type: PDFs with a text layer go to parse_v2(), image-only PDFs are rasterized and go to OCR plus parse_v4(), images whose OCR text contains "INVOICE" are treated as invoice scans, other images get OCR for error codes and request IDs (with a VLM description in live mode), audio gets duration, VAD segments, and a transcription (live) or a transcription quote, and screen recordings are sampled at 1 fps and deduplicated. Every report records the route, fields, status, local time, and any cost estimate.
    • main() loads T-1001 with load_tickets(), processes six attachments (the invoice PDF, the image-only scan, a held-out invoice photo, the error screenshot, the voice note, and the screen recording), prints the reports, and saves them to data/attachments/lab_report.json.
  • Comes out: (local times depend on whether OCR results are already cached from earlier scripts; a cold OCR run takes 1 to 2 seconds per page)

Now run it live and compare:

bash
export GEMINI_API_KEY=your-key-here GROQ_API_KEY=your-other-key
LLM_PROVIDER=gemini M12_LIVE=1 python examples/m12_lab.py

Code explained

  • In simple words: the same router, now allowed to call models.
  • What happens: the failing invoice goes to vlm_extract() with the Gemini vision model and is auto-accepted only if the model's fields pass check_invoice(); the voice note goes to Groq Whisper. The error screenshot's route notes that a VLM description would be added.
  • Comes out: we did not run this for the build, so there is no output to show. Look for the held-out invoice's route to change to image: OCR -> VLM with status auto-accepted after escalation (if the model read the seat count and the arithmetic checks pass) or stay in human review (if not). Either outcome is correct behavior for the router; which one happens is a measurement of the model.

Extend it: (1) Add a screenshot branch that crops the dialog region before OCR and compare times. (2) Add a held-out batch of screenshots with different fonts and sizes and measure the OCR route's accuracy on them. (3) Route the voice note transcript through Module 6's triage schema, so a voice-only ticket gets a category and priority like a typed one.

Project Milestone

After this module, your supportdesk copy contains:

  • requirements.txt with five new pins: pillow==12.3.0, pypdf==6.19.0, pypdfium2==5.13.0, reportlab==5.0.1, rapidocr-onnxruntime==1.4.4.
  • data/attachments/: the generated screenshot, 24 invoices (12 dev, 12 held-out) as PDFs and scans, an image-only PDF, the chart, counting, and fine-print images, the voice note, the screen recording, ground_truth.json, and the OCR cache.
  • examples/m12_make_assets.py: the deterministic attachment generator.
  • examples/m12_image_tokens.py, m12_resize_legibility.py, m12_vision_chat.py, m12_image_qa.py: image token and cost math per provider, resize-versus-legibility measurement, image requests through llm.chat with image-aware token estimates, and an image question harness.
  • examples/m12_pdf_text.py, m12_invoice_pipeline.py, m12_ocr_peek.py, m12_extraction_eval.py: the document pipeline (four parsers, OCR, VLM path, arithmetic checks) and its field-level evaluation on dev and held-out sets.
  • examples/m12_audio.py, m12_video_frames.py, m12_image_edit.py, m12_voice_latency.py: voice notes and transcription, frame sampling, redaction and generation, and the real-time voice budget and turn manager.
  • examples/m12_lab.py: the attachment router, writing data/attachments/lab_report.json.
  • tests/test_m12_multimodal.py: 31 tests.

The assistant can now read what customers attach. It still drafts replies for Maya's team to review, and it never acts on an attachment alone: a parsed invoice is evidence for the refund check from Module 8, not an authorization.

Interview Questions

1. How does a vision-language model "see" an image, and why does that matter for cost? The image is resized to the provider's limits, cut into patches (for example 28 or 32 pixels square), encoded by a vision encoder, and projected into the language model's embedding space as image tokens. Those tokens sit in the context and are billed as input. So cost usually scales with pixel area up to a cap: a 1280x800 screenshot is 1,334 tokens on Claude and 504 if you send it at 768 px. Some providers charge a flat amount instead (Groq's 2,048 per image, Gemini 3's 1,120 default), where resizing saves nothing.

2. A customer's screenshot has a tiny workspace ID in the footer, and the model keeps getting it wrong. What do you do? First check whether the characters survive the provider's resize. In this module, a 14 px footer on a 2400 px screenshot was readable at 1568 px and gone at 1024 px, so a model receiving the downscaled image is guessing. Crop the footer at full resolution and send that (112 tokens, read correctly), or read exact identifiers with OCR and use the VLM only for interpretation. Also add an "unreadable" option to the prompt so the model can decline instead of inventing.

3. Your token budget guard rejected a request with one small screenshot as 44,000 tokens. What happened? The guard counted the base64 data URL as text. Images are billed by their pixel dimensions under the provider's rule, not by the length of their encoding. The fix is an image-aware estimator that counts text parts with the tokenizer and image parts with the provider's formula. Here that gives about 950 to 2,150 tokens for the same request.

4. When would you parse a PDF instead of sending it to a VLM? Whenever it has a text layer and a known layout. The text layer is exact, free, and takes milliseconds: our layout-aware parser read 108 of 108 fields. Check for a text layer first (an empty extraction means a scan), use layout-preserving extraction for multi-column documents, and validate. A VLM earns its cost on scans with unknown layouts, handwriting, or when you need interpretation rather than extraction.

5. Your OCR parser scores 100% on your test invoices. Are you done? Only if those invoices were not the ones you looked at while writing it. Ours scored 100% on the 12 dev invoices and 91.7% (95% CI 85% to 96%) on 12 held-out ones, with perfect documents falling from 12 to 4. The held-out errors were new OCR confusions and digits OCR never detected. Always keep a held-out set, report it, and read its errors before the next iteration.

6. How do you make an imperfect extractor safe to put in production? Add checks that do not need ground truth: required fields present, seats times unit price equals subtotal, subtotal plus tax equals total, formats valid. Auto-accept only what passes; route the rest to a stronger reader or a person. Measure the checks: here they flagged every wrong document on every path with no false alarms. Then sample accepted documents for human review, because a consistent misread can pass arithmetic.

7. What are VLMs bad at, and how do you work around it? Fine detail lost at the resize step, counting many similar or overlapping objects, precise spatial relations (BlindTest: 58% average on circles and line crossings in 2024; BLINK: 51% for GPT-4V against 96% for humans), and exact values from a chart. Workarounds: crop, use OCR or a detector for exact strings and counts, ask for values only when they are printed, allow "unreadable", validate against known constraints, and measure on your own images, because newer models improve unevenly.

8. How would you handle voice messages on support tickets? Transcribe with an ASR model (Groq's whisper-large-v3 costs 0.111 USD per hour, about 4 USD a month for 3,000 45-second notes), then send the transcript through the normal text pipeline. Convert to 16 kHz mono, run VAD first so silence and noise are not transcribed (Whisper can hallucinate phrases over non-speech: about 1% of transcriptions in the "Careless Whisper" study), chunk long recordings with overlap, and keep the audio for verification. Use a speech-native model only when you need tone or non-speech information.

9. How do you estimate the cost of video understanding? Frames times tokens per frame plus seconds times audio tokens per second. On Gemini 3 at default resolution, a minute at 1 fps is 60 x 70 + 60 x 25 = 5,700 tokens; at high resolution it is 18,300. The biggest lever is the sampling rate, and the second is deduplication: our mostly static screen recording kept 11 of 60 frames after dropping near-duplicates, with every screen change still caught.

10. Walk me through the latency budget of a voice agent. Silence endpointing (about 500 ms), streaming ASR (about 150 ms), LLM time to first chunk (0.77 s for gpt-oss-120b on Groq), the first sentence (about 30 ms at 474 tokens per second), TTS time to first audio (about 75 ms), and network hops: about 1.7 s in total, against a human gap of about 200 ms. The biggest items are the model and the endpointing, so use a fast non-reasoning model (high reasoning pushes this to 5.9 s), stream the first sentence to TTS, use a semantic or eager end-of-turn detector (Deepgram reports 150 to 250 ms saved for 50 to 70% more LLM calls), and play filler audio during tool calls.

11. The caller interrupts the agent mid-sentence. What must happen? Stop audio playback immediately, cancel or discard the rest of the generation, and record in the conversation history only what the caller actually heard, marked as interrupted. Otherwise the model will assume the caller heard advice they never got. Also distinguish real interruptions from backchannels ("uh-huh") and background noise with a minimum speech duration or a semantic check.

12. Should you use an image-editing model to redact personal data from screenshots? No. A generative edit redraws pixels and can alter or invent content; redaction must be exact, repeatable, and auditable. Use OCR to find the text, pattern-match what must go, draw opaque rectangles, and verify with a second OCR pass (which found nothing in our test). For known sensitive regions, block the whole region regardless of what OCR detects.

Other Tools and Providers

Tool or providerWhat it isWhen to use it instead of what we did
OpenAI GPT models with vision, Realtime API (gpt-realtime-2.1), GPT Image (gpt-image-2.5-*)Hosted VLMs, speech-to-speech, image generation and editingWhen you are on OpenAI; token rules are in the "Images and vision" guide
Anthropic Claude (vision, PDF input)Hosted VLM with documented 28 px patch pricing and a high-resolution tierDocument and screenshot understanding with predictable image cost
Google Gemini (native API)Images, PDFs, audio, and video in one model, with media_resolution and video fps controlsVideo and long audio; per-part resolution control, which the OpenAI-compatibility page did not list when checked
Ollama vision models (gemma4, qwen3.8)Local VLMsAttachments that must not leave your machine
Tesseract, PaddleOCR, docTR, EasyOCROther OCR enginesTesseract for many languages; the others for accuracy on hard scans (some download models at first run)
Cloud document AI (Google Document AI, AWS Textract, Azure Document Intelligence)Managed OCR, forms, and table extractionHigh-volume forms and tables with managed accuracy and support
pdfplumber, Camelot, Docling, UnstructuredPDF layout, table, and document-conversion librariesComplex tables and mixed documents; conversion to Markdown for RAG
faster-whisper, whisper.cppLocal Whisper inferenceTranscription without sending audio off-device
Deepgram, AssemblyAI, ElevenLabs ScribeHosted streaming ASR with diarization and end-of-turn featuresReal-time voice and call analytics
Silero VAD, WebRTC VADTrained voice activity detectorsReplacing our energy VAD before transcription
LiveKit Agents, PipecatFrameworks for real-time voice agentsTurn-taking, interruption, and transport handled for you
ffmpeg, PySceneDetectVideo decoding and scene-change detectionReal video files (MP4, WebM) instead of our GIF

Model names, image token rules, and prices in this area change every few months. Check the provider pages before committing, and keep the checks in test_m12_multimodal.py pointed at the numbers you rely on.

Coming Up in Module 13

This module ended with per-item costs and latencies: a few thousandths of a dollar per attachment, about a second of CPU per OCR page, 1.7 seconds to a voice reply. Module 13, Deployment, Operations, and Economics, turns those into production decisions: hosted APIs versus self-hosting (including when a local OCR box or a local vision model on your own GPU pays for itself), rate limits and backoff for bursty attachment traffic, fallback chains when a vision model is down, caching, cost-aware routing like our attachment router at larger scale, observability that logs multimodal requests without logging customer images, and how to survive a provider changing its image token rules or retiring the model your pipeline depends on.