Part 4: Multimodal system design
Choosing between a multimodal model and a specialized pipeline
Every attachment type has two families of solution: send it to a general multimodal model, or run a specialized pipeline (parser, OCR, ASR, frame sampler) and send text onward. The measurements in this module point to a consistent rule: use the specialized pipeline for the exact, high-volume, checkable parts, and the multimodal model for interpretation and for whatever the pipeline cannot handle.
| Situation | Use this | Why |
|---|---|---|
| Exact identifiers in screenshots (error codes, request IDs) | OCR plus regex | Measured 6/6 correct on exact strings, local, free, no hallucinated characters |
| "What went wrong in this screenshot?" | VLM | Needs interpretation; no specialized tool does it |
| Born-digital PDFs | Text layer plus parser | 108/108 fields, 2 ms per document |
| Scans of known document types | OCR plus tolerant parser plus checks, VLM on failure | 91.7% of fields held-out, every wrong document flagged |
| Documents of unknown layout | VLM with schema plus checks | One prompt covers many layouts |
| Voice notes | ASR, then the text pipeline | About 4 USD a month at 3,000 notes; transcripts are searchable and redactable |
| Tone, emotion, or non-speech sound | Speech-native model | Transcripts discard it |
| Screen recordings | Sample, deduplicate, then VLM on the kept frames | 11 frames instead of 60 on our recording |
| Counting, measuring, exact positions | A specialized detector, or ask the customer | VLM counting and spatial reasoning are documented weak spots |
| Data that must not leave your infrastructure | Local pipeline, or a local VLM through Ollama | Privacy beats convenience |
Cost and latency of multimodal inputs
Here are this module's numbers side by side. Local times are from the course machine (2 CPU cores) and will differ on yours; token counts are exact under each provider's published rule; dollar figures use prices checked 21 September 2026.
| Input | Reader | Tokens | Cost per item | Local time |
|---|---|---|---|---|
| Born-digital invoice PDF | pypdf plus parser | 0 | 0 | about 2 ms |
| Scanned invoice (one page) | RapidOCR plus parser | 0 | 0 (CPU only) | about 1 s |
| Scanned invoice (one page) | Gemini 3.5 Flash, default resolution | 1,120 image + about 150 text | about 0.0037 USD (with 200 output tokens) | network plus model time |
| Screenshot 1280x800 | Claude Haiku 4.5 | 1,334 (504 at 768 px) | 0.0013 USD (0.0005) | network plus model time |
| Screenshot, any size | Groq Qwen 3.8 27B | 2,048 | 0.0016 USD | network plus model time |
| 45 s voice note | Groq whisper-large-v3 | not token-billed | 0.0014 USD | network plus model time |
| 45 s voice note | Gemini, speech-native | 1,125 to 1,440 | about 0.002 USD at pricing.py's input rate | network plus model time |
| 60 s screen recording, 1 fps | Gemini 3, default resolution | 5,700 including audio | about 0.009 USD | network plus model time |
| 60 s screen recording, deduplicated | 11 frames at Gemini 3 high | about 3,080 | about 0.005 USD | about 90 ms to sample |
Two patterns stand out. Per item, multimodal input is cheap; the costs that bite are volume (every frame, every agent step, every page of a long PDF) and latency (a model call is hundreds of milliseconds to seconds, while a parser takes milliseconds). And the context window fills fast: 60 frames at high resolution is 16,800 image tokens before a word of the prompt, so budgets from Module 2 and Module 7 need image-aware counting, which is what estimate_input_tokens() is for.
Evaluating multimodal outputs
Everything in Module 10 applies; multimodal inputs add a few specifics:
- Generate or label ground truth per field. Synthetic assets (as here) give free, exact labels for mechanics; a few hundred labelled real attachments give the numbers you actually ship on.
- Score fields, not documents only. Per-field accuracy shows which field breaks (v3's
totalandplan); per-document accuracy shows the business impact ("how many invoices need a person?"). Report both, with n and a confidence interval. - Hold out a test set and do not look at it while building. Measured here: 100% on dev, 91.7% held-out, 12 of 12 perfect documents versus 4 of 12.
- Normalize before comparing.
12,480and12480are the same answer;288and288.00are the same amount. Decide the normalization once, in code, and use it for every reader. - Count abstentions separately from errors. A reader that says "unreadable" or
Noneis safer than one that guesses: its failures are routable. v4's held-out misses were all abstentions. - Measure what the checks catch. "Wrong documents flagged" and "false alarms" tell you whether auto-accepting checked outputs is safe.
- Slice by input condition: layout, scan quality, image size, language, audio quality. The average hides the slice that fails.
- For descriptions and summaries, where there is no exact answer, use Module 10's rubric-based LLM-as-judge, and give the judge the image too, or it grades fluency rather than faithfulness.
- Re-run on every model or prompt change. Vision behavior changes between model versions more than text behavior does, and providers change resize rules without changing model names.
Tests for the building blocks
The tests pin down every deterministic piece: token formulas against the providers' examples, the resize caps, both parsers on every generated PDF, the checks, the VLM plumbing with ScriptedLLM, image-aware token estimation, the audio helpers, frame sampling timestamps, barge-in, and type sniffing.
tests/test_m12_multimodal.py
"""Tests for Module 12. Run: PYTHONPATH=. pytest -q tests/test_m12_multimodal.py
OCR-dependent tests are marked slow-ish (a few seconds each on CPU)."""
from __future__ import annotations
import json
import random
import subprocess
import sys
from pathlib import Path
import pytest
from examples.m12_audio import chunk_plan, speech_segments, whisper_cost
from examples.m12_image_tokens import (claude_resize, claude_tokens, gemini_legacy_tokens,
openai_patch_tokens, openai_tile_tokens)
from examples.m12_invoice_pipeline import (FIELDS, KNOWN_ACCOUNTS, best_account, check_invoice, image_message,
money, parse_v1, parse_v2, parse_v4, parse_vlm_json, pdf_text, snap,
vlm_extract)
from examples.m12_voice_latency import STAGES, budget
from examples.m12_video_frames import drop_near_duplicates, frames_at
from examples.m12_vision_chat import estimate_input_tokens, screenshot_messages, encode_image
from examples.m12_voice_latency import TurnManager
from examples.m12_lab import sniff
from supportdesk.stand_in import ScriptedLLM
from supportdesk.tokens import count_messages
ATT = Path("data/attachments")
@pytest.fixture(scope="session", autouse=True)
def assets():
if not (ATT / "ground_truth.json").exists():
subprocess.run([sys.executable, "examples/m12_make_assets.py"], check=True)
return json.loads((ATT / "ground_truth.json").read_text())
@pytest.mark.parametrize("got,want", [
(openai_patch_tokens(1024, 1024), 1229), (openai_patch_tokens(2048, 2048), 3000),
(openai_patch_tokens(4096, 512), 2458), (openai_tile_tokens(1024, 1024), 765),
(openai_tile_tokens(2048, 4096), 1105), (openai_tile_tokens(4000, 3000, detail="low"), 85),
(claude_tokens(200, 200), 64), (claude_tokens(1000, 1000), 1296), (claude_tokens(1092, 1092), 1521),
(claude_tokens(1920, 1080), 1560), (claude_tokens(2000, 1500), 1564), (claude_tokens(3840, 2160), 1560),
(claude_tokens(1920, 1080, "high"), 2691), (claude_tokens(3840, 2160, "high"), 4784),
(gemini_legacy_tokens(960, 540), 1548), (gemini_legacy_tokens(384, 384), 258),
])
def test_token_formulas_match_provider_examples(got, want):
assert got == want
def test_claude_resize_never_enlarges_and_respects_caps():
rng = random.Random(0)
for _ in range(300):
w, h = rng.randint(10, 6000), rng.randint(10, 6000)
rw, rh = claude_resize(w, h)
assert rw <= w and rh <= h and max(rw, rh) <= 1568 and claude_tokens(rw, rh) <= 1568
def test_layout_parser_reads_every_generated_pdf(assets):
for truth in assets["invoices"]:
got = parse_v2(pdf_text(ATT / "invoices" / f"{truth['invoice_number']}.pdf", layout=True))
assert {f: str(got[f]) for f in FIELDS} == {f: str(truth[f]) for f in FIELDS}
def test_naive_parser_fails_on_modern_layout(assets):
modern = next(t for t in assets["invoices"] if t["layout"] == "modern")
got = parse_v1(pdf_text(ATT / "invoices" / f"{modern['invoice_number']}.pdf"))
assert all(v is None for v in got.values())
def test_checks_flag_arithmetic_errors(assets):
good = dict(assets["invoices"][0])
assert check_invoice(good) == []
assert "subtotal + tax != total" in check_invoice({**good, "total": "289.00"})
assert "missing plan" in check_invoice({**good, "plan": None})
def test_snap_repairs_ocr_spacing_but_not_strangers():
assert snap("NorthwindStudio", KNOWN_ACCOUNTS) == "Northwind Studio"
assert snap("AcmeRobotics.", KNOWN_ACCOUNTS) == "Acme Robotics"
assert snap("Totally Unknown LLC", KNOWN_ACCOUNTS) == "Totally Unknown LLC"
def test_vlm_path_plumbing_with_stand_in():
llm = ScriptedLLM(replies=['```json\n{"invoice_number": "INV-2026-004512", "seats": "24", "total": 288}\n```'])
got = vlm_extract(ATT / "invoices" / "INV-2026-004512_scan.png", llm=llm)
parts = llm.calls[0]["messages"][0]["content"]
assert [p["type"] for p in parts] == ["text", "image_url"]
assert parts[1]["image_url"]["url"].startswith("data:image/png;base64,")
assert got["seats"] == 24 and got["total"] == "288.00" and got["customer"] is None
assert parse_vlm_json("no json here") == {f: None for f in FIELDS}
def test_image_messages_are_counted_as_images_not_text():
url, _, size = encode_image(ATT / "error_dialog.png", 1024)
msgs = screenshot_messages(url)
naive = count_messages(msgs)
for provider in ("groq", "gemini", "ollama"):
assert estimate_input_tokens(msgs, provider, [size]) < naive / 10
assert image_message("hi", ATT / "usage_chart.png")[0]["content"][1]["type"] == "image_url"
def test_audio_helpers():
assert len(speech_segments(ATT / "voice_note.wav")) == 6
assert whisper_cost(3, "whisper-large-v3") == whisper_cost(10, "whisper-large-v3")
plan = chunk_plan(2700)
assert plan[0][0] == 0 and plan[-1][1] == 2700
assert all(a[1] > b[0] for a, b in zip(plan, plan[1:])) # consecutive chunks overlap
def test_frame_sampling_and_dedupe(assets):
frames = frames_at(ATT / "screen_recording.gif", fps=1)
assert len(frames) == assets["screen_recording.gif"]["seconds"]
kept = drop_near_duplicates(frames)
assert len(kept) < len(frames)
kept_times = [ts for ts, _ in kept]
assert {12.0, 25.0, 47.0} <= set(kept_times) # every screen change is kept, on time
def test_barge_in_keeps_only_what_was_heard():
tm = TurnManager()
tm.on(0, "user_final", "hi")
tm.on(10, "reply_ready", "one two three four five six")
tm.on(20, "played", "9")
tm.on(30, "user_speech")
assert tm.state == "LISTENING"
assert tm.history[-1]["content"] == "one two [interrupted]"
def test_sniff_uses_bytes_not_names(tmp_path):
fake = tmp_path / "invoice.pdf"
fake.write_bytes((ATT / "error_dialog.png").read_bytes())
assert sniff(fake) == "image"
assert sniff(ATT / "voice_note.wav") == "audio"
assert sniff(ATT / "screen_recording.gif") == "video"
def test_money_takes_last_separator_as_decimal_point():
assert money("Subtotal 552,00") == "552.00"
assert money("Total due:1,320.00 USD") == "1320.00"
assert money("Amiountdue 144:00") == "144.00"
assert money("no amount here") is None
def test_v4_on_known_ocr_noise():
text = ("INVOICE No. 1NV-2026-004549\nBrightanenc panssi: 2026-03-18\nCustomer QuantaLabs'\n"
"Teamplanx10users@12:00 120.00\nSubtotal 120.00\n.VAT/Tax. 24.00.\nAmiountdue 144:00\nCurrency EUR:")
got = parse_v4(text)
assert got == {"invoice_number": "INV-2026-004549", "invoice_date": "2026-03-18", "customer": "Quanta Labs",
"plan": "Team", "seats": 10, "subtotal": "120.00", "tax": "24.00", "total": "144.00",
"currency": "EUR"}
assert check_invoice(got) == []
assert best_account("Totally Unknown LLC", KNOWN_ACCOUNTS) is None
def test_heldout_split_is_separate(assets):
dev = {t["invoice_number"] for t in assets["invoices"]}
held = {t["invoice_number"] for t in assets["invoices_heldout"]}
assert len(dev) == len(held) == 12 and not dev & held
for t in assets["invoices_heldout"]:
assert (ATT / "invoices" / "heldout" / f"{t['invoice_number']}_scan.png").exists()
def test_latency_budget_adds_up():
assert budget(STAGES) == sum(ms for _, ms, _ in STAGES)
assert budget(STAGES, eager_saving_ms=200) == budget(STAGES) - 200
Code explained
- In simple words: 31 fast checks that the arithmetic, parsers, and plumbing still behave after any change.
- What happens: the
assetssession fixture regenerates the attachments ifground_truth.jsonis missing.test_token_formulas_match_provider_examplescompares 16 token computations with the providers' published values;test_claude_resize_never_enlarges_and_respects_capstries 300 random sizes; the parser tests read every generated PDF;test_vlm_path_plumbing_with_stand_inuses ScriptedLLM to check the request shape and type coercion (plumbing, not model quality); the rest cover the checks,money(),parse_v4(), the dev/held-out split, image-aware token estimates, audio helpers, frame timestamps, barge-in, the latency budget, and type sniffing. - Comes out: nothing by itself; run it with the command below.
python -m pytest -q tests/test_m12_multimodal.py
Code explained
- In simple words: run the suite.
- What happens: the session fixture regenerates the attachments if
ground_truth.jsonis missing. The parametrized test compares 16 token computations with the providers' published values. Other tests check that the resize never enlarges and respects caps on 300 random sizes, that v2 reads all 12 dev PDFs, that v1 fails on the modern layout, that the checks flag arithmetic errors, thatmoney()andparse_v4()handle known OCR noise, that dev and held-out do not overlap, that image messages are counted as images, that the frame sampler keeps every screen change at the right second, that barge-in keeps only heard words, and thatsniff()trusts bytes rather than file names. OCR-heavy checks use the cache, so the suite is fast. - Comes out:
Module Lab
Build the attachment router for ticket T-1001 ("Charged twice this month"). Maya's team wants every attachment on a ticket read by the cheapest reader that can be checked, with anything uncertain routed to a person, and the cost of any model escalation shown. The lab combines the type sniffing, PDF parsing, OCR, invoice checks, audio analysis, frame sampling, and cost estimates from this module.
examples/m12_lab.py
"""Module 12 lab: route every attachment on a Brightlane ticket to the cheapest reliable reader.
Real offline: type sniffing, PDF parsing, OCR, arithmetic checks, audio and video
analysis, token and cost estimates. With M12_LIVE=1 and a key, escalations go to a
vision model and voice notes to Whisper; otherwise they go to the human review queue.
Run: PYTHONPATH=. python examples/m12_lab.py
"""
from __future__ import annotations
import json
import os
import re
import time
from pathlib import Path
import pypdfium2 as pdfium
from PIL import Image
from examples.m12_audio import speech_segments, transcribe, wav_info, whisper_cost
from examples.m12_extraction_eval import cached_ocr
from examples.m12_image_tokens import gemini3_tokens
from examples.m12_invoice_pipeline import check_invoice, parse_v2, parse_v4, pdf_text, vlm_extract
from examples.m12_video_frames import drop_near_duplicates, frames_at
from examples.m12_vision_chat import VISION_MODELS
from supportdesk.data import load_tickets
from supportdesk.llm import Usage, resolve
from supportdesk.pricing import cost_usd
ATT = Path("data/attachments")
LIVE = os.environ.get("M12_LIVE") == "1"
VLM_PRICE_MODEL = "gemini-3.5-flash" # used only to estimate what an escalation would cost
def sniff(path: Path) -> str:
"""Decide the type from the first bytes, never from the file name a customer chose."""
head = path.read_bytes()[:12]
if head.startswith(b"%PDF"):
return "pdf"
if head.startswith(b"\x89PNG") or head.startswith(b"\xff\xd8\xff"):
return "image"
if head.startswith(b"GIF8"):
return "video" # we treat animated GIFs as screen recordings
if head[:4] == b"RIFF" and head[8:12] == b"WAVE":
return "audio"
return "unknown"
def vlm_estimate_usd(n_images: int = 1, output_tokens: int = 200) -> float:
return cost_usd(Usage(input_tokens=n_images * gemini3_tokens("high") + 150, output_tokens=output_tokens),
VLM_PRICE_MODEL)
def handle_invoice_fields(fields: dict, source: Path, how: str) -> dict:
issues = check_invoice(fields)
if not issues:
return {"route": how, "fields": fields, "status": "auto-accepted"}
if LIVE:
provider, _ = resolve()
image = source if source.suffix == ".png" else rasterize(source)
vlm_fields = vlm_extract(image, model=VISION_MODELS[provider])
if not check_invoice(vlm_fields):
return {"route": how + " -> VLM", "fields": vlm_fields, "status": "auto-accepted after escalation"}
return {"route": how, "fields": fields, "status": f"human review ({'; '.join(issues)})",
"escalation_estimate_usd": round(vlm_estimate_usd(), 5)}
def rasterize(pdf: Path, dpi: int = 100) -> Path:
out = pdf.with_name(pdf.stem + "_page1.png")
pdfium.PdfDocument(str(pdf))[0].render(scale=dpi / 72).to_pil().save(out)
return out
def process(path: Path) -> dict:
kind = sniff(path)
started = time.perf_counter()
if kind == "pdf":
text = pdf_text(path, layout=True)
if len(text.strip()) > 50: # a real text layer: parse it, no model needed
report = handle_invoice_fields(parse_v2(text), path, "pdf text layer")
else: # image-only PDF: rasterize, OCR, tolerant parser
report = handle_invoice_fields(parse_v4(cached_ocr(rasterize(path))), path, "scanned pdf: OCR")
elif kind == "image" and "INVOICE" in cached_ocr(path).upper(): # a photographed or scanned invoice
report = handle_invoice_fields(parse_v4(cached_ocr(path)), path, "image: OCR")
elif kind == "image":
text = cached_ocr(path) # cheap local pass for exact identifiers
codes = sorted(set(re.findall(r"\bE-\d{4}\b", text)))
ids = re.findall(r"Request\s*ID:?\s*([0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4})", text)
report = {"route": "OCR for ids" + (" + VLM for description" if LIVE else ""),
"fields": {"error_codes": codes, "request_ids": ids},
"status": "ids extracted" if codes else "human review (no error code found)",
"vlm_description_estimate_usd": round(vlm_estimate_usd(), 5)}
elif kind == "audio":
info = wav_info(path)
report = {"route": "speech-to-text", "fields": {"seconds": round(info["seconds"], 1),
"speech_segments": len(speech_segments(path))},
"status": "transcribed" if LIVE else "queued for transcription",
"transcription_usd": round(whisper_cost(info["seconds"], "whisper-large-v3"), 6)}
if LIVE:
report["fields"]["transcript"] = transcribe(path)["text"]
elif kind == "video":
kept = drop_near_duplicates(frames_at(path, fps=1))
report = {"route": "1 fps sampling + near-duplicate drop",
"fields": {"frames_kept": len(kept), "at_seconds": [ts for ts, _ in kept]},
"status": "frames ready for VLM",
"vlm_estimate_usd": round(vlm_estimate_usd(n_images=len(kept)), 5)}
else:
report = {"route": "rejected", "fields": {}, "status": "unsupported type"}
report.update(file=str(path.relative_to(ATT)), kind=kind,
local_ms=round((time.perf_counter() - started) * 1000))
return report
def main() -> None:
ticket = next(t for t in load_tickets() if t.id == "T-1001")
attachments = [ATT / "invoices/INV-2026-004512.pdf", ATT / "invoices/INV-2026-004514_scanned.pdf",
ATT / "invoices/heldout/INV-2026-005528_scan.png", ATT / "error_dialog.png",
ATT / "voice_note.wav", ATT / "screen_recording.gif"]
print(f"{ticket.id}: {ticket.subject!r} with {len(attachments)} attachments (live={LIVE})\n")
reports = [process(p) for p in attachments]
for r in reports:
print(f"{r['file']} [{r['kind']}] via {r['route']} ({r['local_ms']} ms local)")
print(f" status: {r['status']}")
print(f" fields: {json.dumps(r['fields'])}")
extra = {k: v for k, v in r.items() if k.endswith("_usd")}
if extra:
print(f" cost: {extra}")
auto = sum(r["status"].startswith(("auto", "ids", "frames", "transcribed")) for r in reports)
spend = sum(v for r in reports for k, v in r.items() if k.endswith("_usd"))
print(f"\nhandled without a person: {auto}/{len(reports)}; model spend if every estimate were used: ${spend:.4f}")
Path("data/attachments/lab_report.json").write_text(json.dumps(reports, indent=2))
if __name__ == "__main__":
main()
Code explained
- In simple words: a mail room for attachments: look at what each file really is, send it to the right desk, check the desk's work, and only pay a model (or a person) when the checks fail.
- What happens:
sniff()decides the type from the file's first bytes (magic numbers:%PDF, the PNG signature,GIF8,RIFF....WAVE), never from the file name or extension a customer chose. Module 11's rule: attachments are untrusted input.handle_invoice_fields()runscheck_invoice(). Clean results are auto-accepted. Otherwise, in live mode it escalates to a VLM and accepts that result only if it passes the same checks; offline, it routes to human review with the reasons and the estimated cost of escalating.rasterize()renders an image-only PDF page with pypdfium2 so OCR (or a VLM) can read it.process()routes each type: PDFs with a text layer go toparse_v2(), image-only PDFs are rasterized and go to OCR plusparse_v4(), images whose OCR text contains "INVOICE" are treated as invoice scans, other images get OCR for error codes and request IDs (with a VLM description in live mode), audio gets duration, VAD segments, and a transcription (live) or a transcription quote, and screen recordings are sampled at 1 fps and deduplicated. Every report records the route, fields, status, local time, and any cost estimate.main()loads T-1001 withload_tickets(), processes six attachments (the invoice PDF, the image-only scan, a held-out invoice photo, the error screenshot, the voice note, and the screen recording), prints the reports, and saves them todata/attachments/lab_report.json.
- Comes out: (local times depend on whether OCR results are already cached from earlier scripts; a cold OCR run takes 1 to 2 seconds per page)
Now run it live and compare:
export GEMINI_API_KEY=your-key-here GROQ_API_KEY=your-other-key
LLM_PROVIDER=gemini M12_LIVE=1 python examples/m12_lab.py
Code explained
- In simple words: the same router, now allowed to call models.
- What happens: the failing invoice goes to
vlm_extract()with the Gemini vision model and is auto-accepted only if the model's fields passcheck_invoice(); the voice note goes to Groq Whisper. The error screenshot's route notes that a VLM description would be added. - Comes out: we did not run this for the build, so there is no output to show. Look for the held-out invoice's route to change to
image: OCR -> VLMwith statusauto-accepted after escalation(if the model read the seat count and the arithmetic checks pass) or stay in human review (if not). Either outcome is correct behavior for the router; which one happens is a measurement of the model.
Extend it: (1) Add a screenshot branch that crops the dialog region before OCR and compare times. (2) Add a held-out batch of screenshots with different fonts and sizes and measure the OCR route's accuracy on them. (3) Route the voice note transcript through Module 6's triage schema, so a voice-only ticket gets a category and priority like a typed one.
Project Milestone
After this module, your supportdesk copy contains:
requirements.txtwith five new pins:pillow==12.3.0,pypdf==6.19.0,pypdfium2==5.13.0,reportlab==5.0.1,rapidocr-onnxruntime==1.4.4.data/attachments/: the generated screenshot, 24 invoices (12 dev, 12 held-out) as PDFs and scans, an image-only PDF, the chart, counting, and fine-print images, the voice note, the screen recording,ground_truth.json, and the OCR cache.examples/m12_make_assets.py: the deterministic attachment generator.examples/m12_image_tokens.py,m12_resize_legibility.py,m12_vision_chat.py,m12_image_qa.py: image token and cost math per provider, resize-versus-legibility measurement, image requests throughllm.chatwith image-aware token estimates, and an image question harness.examples/m12_pdf_text.py,m12_invoice_pipeline.py,m12_ocr_peek.py,m12_extraction_eval.py: the document pipeline (four parsers, OCR, VLM path, arithmetic checks) and its field-level evaluation on dev and held-out sets.examples/m12_audio.py,m12_video_frames.py,m12_image_edit.py,m12_voice_latency.py: voice notes and transcription, frame sampling, redaction and generation, and the real-time voice budget and turn manager.examples/m12_lab.py: the attachment router, writingdata/attachments/lab_report.json.tests/test_m12_multimodal.py: 31 tests.
The assistant can now read what customers attach. It still drafts replies for Maya's team to review, and it never acts on an attachment alone: a parsed invoice is evidence for the refund check from Module 8, not an authorization.
Interview Questions
1. How does a vision-language model "see" an image, and why does that matter for cost? The image is resized to the provider's limits, cut into patches (for example 28 or 32 pixels square), encoded by a vision encoder, and projected into the language model's embedding space as image tokens. Those tokens sit in the context and are billed as input. So cost usually scales with pixel area up to a cap: a 1280x800 screenshot is 1,334 tokens on Claude and 504 if you send it at 768 px. Some providers charge a flat amount instead (Groq's 2,048 per image, Gemini 3's 1,120 default), where resizing saves nothing.
2. A customer's screenshot has a tiny workspace ID in the footer, and the model keeps getting it wrong. What do you do? First check whether the characters survive the provider's resize. In this module, a 14 px footer on a 2400 px screenshot was readable at 1568 px and gone at 1024 px, so a model receiving the downscaled image is guessing. Crop the footer at full resolution and send that (112 tokens, read correctly), or read exact identifiers with OCR and use the VLM only for interpretation. Also add an "unreadable" option to the prompt so the model can decline instead of inventing.
3. Your token budget guard rejected a request with one small screenshot as 44,000 tokens. What happened? The guard counted the base64 data URL as text. Images are billed by their pixel dimensions under the provider's rule, not by the length of their encoding. The fix is an image-aware estimator that counts text parts with the tokenizer and image parts with the provider's formula. Here that gives about 950 to 2,150 tokens for the same request.
4. When would you parse a PDF instead of sending it to a VLM? Whenever it has a text layer and a known layout. The text layer is exact, free, and takes milliseconds: our layout-aware parser read 108 of 108 fields. Check for a text layer first (an empty extraction means a scan), use layout-preserving extraction for multi-column documents, and validate. A VLM earns its cost on scans with unknown layouts, handwriting, or when you need interpretation rather than extraction.
5. Your OCR parser scores 100% on your test invoices. Are you done? Only if those invoices were not the ones you looked at while writing it. Ours scored 100% on the 12 dev invoices and 91.7% (95% CI 85% to 96%) on 12 held-out ones, with perfect documents falling from 12 to 4. The held-out errors were new OCR confusions and digits OCR never detected. Always keep a held-out set, report it, and read its errors before the next iteration.
6. How do you make an imperfect extractor safe to put in production? Add checks that do not need ground truth: required fields present, seats times unit price equals subtotal, subtotal plus tax equals total, formats valid. Auto-accept only what passes; route the rest to a stronger reader or a person. Measure the checks: here they flagged every wrong document on every path with no false alarms. Then sample accepted documents for human review, because a consistent misread can pass arithmetic.
7. What are VLMs bad at, and how do you work around it? Fine detail lost at the resize step, counting many similar or overlapping objects, precise spatial relations (BlindTest: 58% average on circles and line crossings in 2024; BLINK: 51% for GPT-4V against 96% for humans), and exact values from a chart. Workarounds: crop, use OCR or a detector for exact strings and counts, ask for values only when they are printed, allow "unreadable", validate against known constraints, and measure on your own images, because newer models improve unevenly.
8. How would you handle voice messages on support tickets? Transcribe with an ASR model (Groq's whisper-large-v3 costs 0.111 USD per hour, about 4 USD a month for 3,000 45-second notes), then send the transcript through the normal text pipeline. Convert to 16 kHz mono, run VAD first so silence and noise are not transcribed (Whisper can hallucinate phrases over non-speech: about 1% of transcriptions in the "Careless Whisper" study), chunk long recordings with overlap, and keep the audio for verification. Use a speech-native model only when you need tone or non-speech information.
9. How do you estimate the cost of video understanding? Frames times tokens per frame plus seconds times audio tokens per second. On Gemini 3 at default resolution, a minute at 1 fps is 60 x 70 + 60 x 25 = 5,700 tokens; at high resolution it is 18,300. The biggest lever is the sampling rate, and the second is deduplication: our mostly static screen recording kept 11 of 60 frames after dropping near-duplicates, with every screen change still caught.
10. Walk me through the latency budget of a voice agent. Silence endpointing (about 500 ms), streaming ASR (about 150 ms), LLM time to first chunk (0.77 s for gpt-oss-120b on Groq), the first sentence (about 30 ms at 474 tokens per second), TTS time to first audio (about 75 ms), and network hops: about 1.7 s in total, against a human gap of about 200 ms. The biggest items are the model and the endpointing, so use a fast non-reasoning model (high reasoning pushes this to 5.9 s), stream the first sentence to TTS, use a semantic or eager end-of-turn detector (Deepgram reports 150 to 250 ms saved for 50 to 70% more LLM calls), and play filler audio during tool calls.
11. The caller interrupts the agent mid-sentence. What must happen? Stop audio playback immediately, cancel or discard the rest of the generation, and record in the conversation history only what the caller actually heard, marked as interrupted. Otherwise the model will assume the caller heard advice they never got. Also distinguish real interruptions from backchannels ("uh-huh") and background noise with a minimum speech duration or a semantic check.
12. Should you use an image-editing model to redact personal data from screenshots? No. A generative edit redraws pixels and can alter or invent content; redaction must be exact, repeatable, and auditable. Use OCR to find the text, pattern-match what must go, draw opaque rectangles, and verify with a second OCR pass (which found nothing in our test). For known sensitive regions, block the whole region regardless of what OCR detects.
Other Tools and Providers
| Tool or provider | What it is | When to use it instead of what we did |
|---|---|---|
OpenAI GPT models with vision, Realtime API (gpt-realtime-2.1), GPT Image (gpt-image-2.5-*) | Hosted VLMs, speech-to-speech, image generation and editing | When you are on OpenAI; token rules are in the "Images and vision" guide |
| Anthropic Claude (vision, PDF input) | Hosted VLM with documented 28 px patch pricing and a high-resolution tier | Document and screenshot understanding with predictable image cost |
| Google Gemini (native API) | Images, PDFs, audio, and video in one model, with media_resolution and video fps controls | Video and long audio; per-part resolution control, which the OpenAI-compatibility page did not list when checked |
Ollama vision models (gemma4, qwen3.8) | Local VLMs | Attachments that must not leave your machine |
| Tesseract, PaddleOCR, docTR, EasyOCR | Other OCR engines | Tesseract for many languages; the others for accuracy on hard scans (some download models at first run) |
| Cloud document AI (Google Document AI, AWS Textract, Azure Document Intelligence) | Managed OCR, forms, and table extraction | High-volume forms and tables with managed accuracy and support |
| pdfplumber, Camelot, Docling, Unstructured | PDF layout, table, and document-conversion libraries | Complex tables and mixed documents; conversion to Markdown for RAG |
| faster-whisper, whisper.cpp | Local Whisper inference | Transcription without sending audio off-device |
| Deepgram, AssemblyAI, ElevenLabs Scribe | Hosted streaming ASR with diarization and end-of-turn features | Real-time voice and call analytics |
| Silero VAD, WebRTC VAD | Trained voice activity detectors | Replacing our energy VAD before transcription |
| LiveKit Agents, Pipecat | Frameworks for real-time voice agents | Turn-taking, interruption, and transport handled for you |
| ffmpeg, PySceneDetect | Video decoding and scene-change detection | Real video files (MP4, WebM) instead of our GIF |
Model names, image token rules, and prices in this area change every few months. Check the provider pages before committing, and keep the checks in test_m12_multimodal.py pointed at the numbers you rely on.
Coming Up in Module 13
This module ended with per-item costs and latencies: a few thousandths of a dollar per attachment, about a second of CPU per OCR page, 1.7 seconds to a voice reply. Module 13, Deployment, Operations, and Economics, turns those into production decisions: hosted APIs versus self-hosting (including when a local OCR box or a local vision model on your own GPU pays for itself), rate limits and backoff for bursty attachment traffic, fallback chains when a vision model is down, caching, cost-aware routing like our attachment router at larger scale, observability that logs multimodal requests without logging customer images, and how to survive a provider changing its image token rules or retiring the model your pipeline depends on.