Part 2: Documents
Invoices are the most common document attachment at Brightlane: "I was charged twice, here is the invoice." The assistant needs nine fields from each one (number, date, customer, plan, seats, subtotal, tax, total, currency) to check the claim against the billing system. This part builds three ways to get them and measures each against the answer key.
PDFs: born-digital versus scanned
"PDF" covers two very different things:
- A born-digital PDF (also called a text-layer PDF) was written by software, such as Brightlane's invoicing system. Besides drawing the page, it stores the characters and their positions. You can extract that text exactly, with no model and no OCR, in milliseconds.
- A scanned PDF (an image-only PDF) contains a picture of a page. There is no text inside it at all. To read it you need OCR or a vision model.
Customers send both, often with the same file name, so the first step is always to check which one you have. Here is what pypdf sees in each:
"""What does a PDF's text layer look like? Plain versus layout extraction on both invoice layouts.
Run: PYTHONPATH=. python examples/m12_pdf_text.py
"""
from __future__ import annotations
from pathlib import Path
from pypdf import PdfReader
INVOICES = Path("data/attachments/invoices")
for name in ("INV-2026-004507.pdf", "INV-2026-004514_scanned.pdf"):
page = PdfReader(INVOICES / name).pages[0]
plain = page.extract_text()
print(f"=== {name}: plain extraction ({len(plain)} characters)")
print(plain)
if plain.strip():
print(f"=== {name}: layout extraction, non-empty lines")
print("\n".join(line.rstrip() for line in page.extract_text(extraction_mode="layout").splitlines() if line.strip()))
Code explained
- In simple words: ask two PDFs for their text; one has plenty, the other has none.
- What happens:
PdfReader(...).pages[0].extract_text()returns the text layer in the order the PDF draws it.extraction_mode="layout"(available in pypdf 3.x and later; we use 6.19.0) instead places each text run at its position on the page, which rebuilds columns and rows. The script prints both for the modern-layout invoice and tries the scanned one. - Comes out:text
=== INV-2026-004507.pdf: plain extraction (227 characters) Subtotal 624.00 VAT / Tax 124.80 Amount due 748.80 Currency EUR INVOICE Brightlane Inc. / 500 Market Street, San Francisco No. INV-2026-004507 Issued 2026-05-21 Customer Acme Robotics Business plan x 26 users @ 24.00 624.00 === INV-2026-004507.pdf: layout extraction, non-empty lines INVOICE No. INV-2026-004507 Brightlane Inc. / 500 Market Street, San Francisco Issued 2026-05-21 Customer Acme Robotics Business plan x 26 users @ 24.00 624.00 Subtotal 624.00 VAT / Tax 124.80 Amount due 748.80 Currency EUR === INV-2026-004514_scanned.pdf: plain extraction (0 characters)The plain extraction of the modern invoice starts with "Subtotal": that layout draws the totals block before the header, and plain extraction follows drawing order, not reading order. Labels and values also come out on separate lines, because they are separate text runs. Layout mode puts every label and its value back on one line. The scanned PDF returns zero characters. A length check on extracted text is therefore the cheapest scanned-document detector there is, and our lab uses it (fewer than 50 characters means "treat as an image").
Tables, forms, and layout
An invoice is a form (labelled fields: "Invoice number: ...") plus a table (line items with columns for description, quantity, unit price, amount). Both are defined by layout: the meaning of "24" comes from the column it sits in. Text extraction flattens layout into a string, and every extraction tool has to decide how:
| Situation | Use this | Why |
|---|---|---|
| Born-digital PDF, one known template | Plain text extraction plus label regexes | Fast and exact; breaks when the template changes (measured below: 0 of 9 fields on the other layout) |
| Born-digital PDF, several templates or multi-column | Layout-preserving extraction (pypdf layout mode, or pdfplumber word positions) | Keeps labels and values on one line; measured: 108 of 108 fields |
| Real tables with ruling lines, spanning cells | A table extractor (pdfplumber, Camelot) or a document-AI service | Reconstructs cells from lines and positions |
| Scans and photos | OCR with a reading-order step, or a VLM | There is no text layer to extract |
| Unknown layouts from many vendors | A VLM (or document-AI service) with a schema, plus validation | Parsers need a template per layout; VLMs generalize |
The parsing pipeline
Here is the whole invoice module: two text-layer parsers, OCR for scans with two OCR-tolerant parsers, the VLM path, and arithmetic cross-checks that work on any path's output. It is long; the Code explained box below walks through every function, and the sections after it measure each piece.
examples/m12_invoice_pipeline.py
"""Three ways to read a Brightlane invoice: PDF text layer, OCR on a scan, and a VLM.
Paths 1 and 2 run for real offline. Path 3 needs an API key; without one it runs
through ScriptedLLM, which tests the plumbing only (it is not a model).
Run: PYTHONPATH=. python examples/m12_invoice_pipeline.py
"""
from __future__ import annotations
import base64
import difflib
import json
import re
from pathlib import Path
from pypdf import PdfReader
from rapidocr_onnxruntime import RapidOCR
from supportdesk.llm import chat
from supportdesk.stand_in import ScriptedLLM
INVOICES = Path("data/attachments/invoices")
FIELDS = ("invoice_number", "invoice_date", "customer", "plan", "seats",
"subtotal", "tax", "total", "currency")
_OCR = None
# ---------- Path 1: the PDF's own text layer ----------
def pdf_text(path: Path, layout: bool = False) -> str:
"""Text of page 1. layout=True keeps the visual arrangement (columns, rows)."""
page = PdfReader(path).pages[0]
return page.extract_text(extraction_mode="layout") if layout else page.extract_text()
def parse_v1(text: str) -> dict:
"""First attempt: 'Label: value' patterns written while looking at one invoice."""
patterns = {
"invoice_number": r"Invoice number: (\S+)",
"invoice_date": r"Invoice date: (\S+)",
"customer": r"Bill to: (.+)",
"plan": r"Brightlane (\w+) plan",
"seats": r"per user\)\n(\d+)",
"subtotal": r"Subtotal: ([\d.]+)",
"tax": r"Tax: ([\d.]+)",
"total": r"Total due: ([\d.]+)",
"currency": r"Total due: [\d.]+ ([A-Z]{3})",
}
out = {}
for field, pattern in patterns.items():
m = re.search(pattern, text)
out[field] = m.group(1).strip() if m else None
if out["seats"] is not None:
out["seats"] = int(out["seats"])
return out
MONEY = r"([\d,]+\.\d{2})"
def parse_v2(text: str) -> dict:
"""Second attempt: both layouts, label synonyms, flexible whitespace. Expects layout text."""
patterns = {
"invoice_number": r"(INV-\d{4}-\d{6})",
"invoice_date": r"(?:Invoice date:?|Issued)\s+(\d{4}-\d{2}-\d{2})",
"customer": r"(?:Bill to:?|Customer)[ \t]+(\S.*?)[ \t]*$",
"plan": r"\b(Team|Business|Enterprise) plan",
"seats": r"(?:per user\)\s+(\d+)\s)|(?:x\s+(\d+)\s+users)",
"subtotal": r"Subtotal:?\s+" + MONEY,
"tax": r"(?:VAT / Tax|Tax):?\s+" + MONEY,
"total": r"(?:Total due|Amount due):?\s+" + MONEY,
"currency": r"(?:(?:Total due|Amount due):?\s+[\d,.]+\s+([A-Z]{3}))|(?:Currency\s+([A-Z]{3}))",
}
out = {}
for field, pattern in patterns.items():
m = re.search(pattern, text, flags=re.MULTILINE)
value = next((g for g in m.groups() if g), None) if m else None
out[field] = value.replace(",", "").strip() if value else None
if out["seats"] is not None:
out["seats"] = int(out["seats"])
return out
# ---------- Path 2: OCR on the scanned image ----------
def ocr_lines(image_path: Path) -> str:
"""OCR, then rebuild reading order: group boxes into rows by height, sort each row left to right."""
global _OCR
if _OCR is None:
_OCR = RapidOCR(intra_op_num_threads=1, inter_op_num_threads=1)
result, _ = _OCR(str(image_path))
boxes = []
for box, text, _score in result or []:
ys = [p[1] for p in box]
boxes.append((sum(ys) / 4, min(p[0] for p in box), max(ys) - min(ys), text))
boxes.sort()
rows: list[list[tuple]] = []
for yc, x, height, text in boxes:
if rows and abs(rows[-1][0][0] - yc) < 0.6 * height:
rows[-1].append((yc, x, height, text))
else:
rows.append([(yc, x, height, text)])
return "\n".join(" ".join(t for *_, t in sorted(row, key=lambda b: b[1])) for row in rows)
KNOWN_ACCOUNTS = ["Northwind Studio", "Acme Robotics", "Blue Harbor Legal", "Kite & Key Design",
"Mendez Logistics", "Orchid Health", "Pine Street Books", "Quanta Labs",
"Riverstone Farms", "Solaris Media", "Tidewater Clinic", "Umbra Games"] # stands in for the CRM
def label(words: str) -> str:
"""Regex for a label that survives OCR: spaces optional, l/I/1/! confusable, stray punctuation."""
parts = []
for ch in words:
if ch == " ":
parts.append(r"\s*")
elif ch in "lI1":
parts.append(r"[lI1!|]")
else:
parts.append(re.escape(ch))
return "".join(parts) + r"\s*[:.;,!]?\s*"
OCR_MONEY = r"(\d[\d,]*[.:]\d{2})" # OCR sometimes reads the decimal point as a colon
def parse_v3(text: str) -> dict:
"""OCR-tolerant parser: fuzzy labels, repaired numbers, customer snapped to a known account."""
total = f"(?:{label('Total due')}|{label('Amount due')})"
patterns = {
"invoice_number": r"(INV-\d{4}-\d{6})",
"invoice_date": f"(?:{label('Invoice date')}|{label('Issued')})" + r"(\d{4}-\d{2}-\d{2})",
"customer": f"(?:{label('Bill to')}|{label('Customer')})" + r"(\S.*?)(?:\s{4,}|$)",
"plan": r"\b(Team|Business|Enterprise)\s*plan",
"seats": r"(?:per\s*user\)\s+(\d+)\s)|(?:x\s*(\d+)\s*users)",
"subtotal": label("Subtotal") + OCR_MONEY,
"tax": f"(?:{label('VAT / Tax')}|{label('Tax')})" + OCR_MONEY,
"total": total + OCR_MONEY,
"currency": f"(?:{total}" + r"[\d,.:]+\s*([A-Z]{3}))|(?:" + label("Currency") + r"([A-Z]{3}))",
}
out = {}
for field, pattern in patterns.items():
m = re.search(pattern, text, flags=re.MULTILINE)
value = next((g for g in m.groups() if g), None) if m else None
out[field] = value.strip() if value else None
for money in ("subtotal", "tax", "total"):
if out[money]:
out[money] = out[money].replace(",", "").replace(":", ".")
if out["seats"] is not None:
out["seats"] = int(out["seats"])
if out["customer"]:
out["customer"] = snap(out["customer"], KNOWN_ACCOUNTS)
return out
def snap(value: str, known: list[str], cutoff: float = 0.8) -> str:
"""Replace a noisy string with the closest known value, if one is close enough."""
squash = {re.sub(r"[^a-z0-9]", "", k.lower()): k for k in known}
hit = difflib.get_close_matches(re.sub(r"[^a-z0-9]", "", value.lower()), list(squash), n=1, cutoff=cutoff)
return squash[hit[0]] if hit else value
# Labels as letters only (OCR mangles spaces and punctuation, so we compare squashed keys).
LABELS = {"invoicedate": "invoice_date", "issued": "invoice_date", "billto": "customer",
"customer": "customer", "subtotal": "subtotal", "tax": "tax", "vattax": "tax",
"totaldue": "total", "amountdue": "total"}
def money(text: str) -> str | None:
"""First amount like 288.00, 1,320.00, 552,00 or 144:00; the last separator is the decimal point."""
m = re.search(r"\d[\d,.:]*[.,:]\d{2}(?!\d)", text)
if not m:
return None
digits = re.sub(r"[^\d]", "", m.group(0))
return f"{digits[:-2]}.{digits[-2:]}"
def best_account(text: str, known: list[str], cutoff: float = 0.8) -> str | None:
"""Which known account name appears in the text? Fuzzy window match on squashed letters."""
squashed = re.sub(r"[^a-z]", "", text.lower())
best, best_score = None, cutoff
for name in known:
key = re.sub(r"[^a-z]", "", name.lower())
for start in range(0, max(1, len(squashed) - len(key) + 1)):
score = difflib.SequenceMatcher(None, key, squashed[start:start + len(key)]).ratio()
if score > best_score:
best, best_score = name, score
return best
def parse_v4(text: str) -> dict:
"""Line-based parser: fuzzy label keys instead of exact label regexes, plus document-wide fallbacks.
Written after reading the v3 errors on the 12 dev invoices. Test it on invoices it has never seen.
"""
out: dict = {f: None for f in FIELDS}
for line in text.splitlines():
head = re.split(r"\d", line, maxsplit=1)[0] # everything before the first digit
key = re.sub(r"[^a-z]", "", head.lower())
match = difflib.get_close_matches(key, list(LABELS), n=1, cutoff=0.75)
if not match:
continue
field = LABELS[match[0]]
value = line[len(head):]
if field in ("subtotal", "tax", "total") and out[field] is None:
out[field] = money(value)
elif field == "invoice_date" and out[field] is None:
m = re.search(r"\d{4}-\d{2}-\d{2}", value)
out[field] = m.group(0) if m else None
number = re.search(r"[I1l|]NV-(\d{4}-\d{6})", text)
out["invoice_number"] = f"INV-{number.group(1)}" if number else None
if out["invoice_date"] is None: # fallback: the only ISO date on the page
m = re.search(r"20\d{2}-\d{2}-\d{2}", text)
out["invoice_date"] = m.group(0) if m else None
plan = re.search(r"(Team|Business|Enterprise)[^A-Za-z\n]{0,3}plan", text)
out["plan"] = plan.group(1) if plan else None
seats = (re.search(r"x\s*(\d+)\s*users", text)
or re.search(r"plan.*?\)\s+(\d+)\s+\d+[.,:]\d{2}", text)) # classic: Qty column after description
out["seats"] = int(seats.group(1)) if seats else None
currency = re.search(r"(?<![A-Z])(USD|EUR|GBP)(?![A-Z])", text)
out["currency"] = currency.group(1) if currency else None
out["customer"] = best_account(text, KNOWN_ACCOUNTS)
return out
# ---------- Path 3: a vision-language model ----------
VLM_PROMPT = (
"You are reading a Brightlane invoice attached to a support ticket. Return only JSON with keys "
+ ", ".join(FIELDS) + ". Use null for any field you cannot read; never guess. "
"Dates as YYYY-MM-DD, money as plain decimals like 288.00, seats as an integer."
)
def image_message(prompt: str, image_path: Path) -> list[dict]:
"""An OpenAI-compatible user message with text plus one base64 data URL image."""
mime = {".png": "image/png", ".jpg": "image/jpeg", ".jpeg": "image/jpeg", ".webp": "image/webp"}[image_path.suffix]
b64 = base64.b64encode(image_path.read_bytes()).decode("ascii")
return [{"role": "user", "content": [
{"type": "text", "text": prompt},
{"type": "image_url", "image_url": {"url": f"data:{mime};base64,{b64}"}},
]}]
def parse_vlm_json(text: str) -> dict:
"""Take the first {...} block; tolerate code fences; coerce types like the other paths."""
m = re.search(r"\{.*\}", text, flags=re.DOTALL)
data = json.loads(m.group(0)) if m else {}
out = {f: data.get(f) for f in FIELDS}
for money in ("subtotal", "tax", "total"):
if isinstance(out[money], (int, float)):
out[money] = f"{out[money]:.2f}"
if isinstance(out["seats"], str) and out["seats"].isdigit():
out["seats"] = int(out["seats"])
return out
def vlm_extract(image_path: Path, llm=chat, **kwargs) -> dict:
result = llm(image_message(VLM_PROMPT, image_path), temperature=0.0, **kwargs)
return parse_vlm_json(result.text)
# ---------- Cross-checks that catch misreads on any path ----------
def check_invoice(fields: dict) -> list[str]:
"""Arithmetic and format checks. They cannot prove a value right, but they catch many wrong ones."""
issues = [f"missing {f}" for f in FIELDS if fields.get(f) in (None, "")]
try:
unit = {"Team": 12.0, "Business": 24.0}[fields["plan"]]
if abs(fields["seats"] * unit - float(fields["subtotal"])) > 0.005:
issues.append("seats x unit price != subtotal")
if abs(float(fields["subtotal"]) + float(fields["tax"]) - float(fields["total"])) > 0.005:
issues.append("subtotal + tax != total")
except (KeyError, TypeError, ValueError):
issues.append("could not run arithmetic checks")
if fields.get("invoice_date") and not re.fullmatch(r"2026-\d{2}-\d{2}", fields["invoice_date"]):
issues.append("bad date format")
return issues
def main() -> None:
truth = json.loads(Path("data/attachments/ground_truth.json").read_text())["invoices"]
for inv in (truth[0], truth[1]):
pdf = INVOICES / f"{inv['invoice_number']}.pdf"
print(f"== {pdf.name} ({inv['layout']} layout)")
v1 = parse_v1(pdf_text(pdf))
v2 = parse_v2(pdf_text(pdf, layout=True))
ocr_text = ocr_lines(pdf.with_name(pdf.stem + "_scan.png"))
ocr2, ocr3, ocr4 = parse_v2(ocr_text), parse_v3(ocr_text), parse_v4(ocr_text)
for name, got in [("pdf v1", v1), ("pdf v2", v2), ("ocr v2", ocr2), ("ocr v3", ocr3), ("ocr v4", ocr4)]:
wrong = [f for f in FIELDS if str(got[f]) != str(inv[f])]
print(f" {name}: {len(FIELDS) - len(wrong)}/{len(FIELDS)} fields right"
+ (f"; wrong or missing: {', '.join(wrong)}" if wrong else "")
+ f"; checks: {check_invoice(got) or 'pass'}")
# VLM path, plumbing only: ScriptedLLM plays back a reply shaped like a model's.
scan = INVOICES / "INV-2026-004512_scan.png"
stand_in = ScriptedLLM(replies=['```json\n{"invoice_number": "INV-2026-004512", "invoice_date": "2026-09-03", '
'"customer": "Northwind Studio", "plan": "Team", "seats": 24, "subtotal": 288, '
'"tax": 0, "total": 288.0, "currency": "USD"}\n```'])
fields = vlm_extract(scan, llm=stand_in)
sent = stand_in.calls[0]["messages"][0]["content"]
print(f"\nVLM path (ScriptedLLM, not a model): parts sent = {[p['type'] for p in sent]}, "
f"data URL length = {len(sent[1]['image_url']['url']):,} chars")
print(f" parsed: {fields}")
print(f" checks: {check_invoice(fields) or 'pass'}")
if __name__ == "__main__":
main()
Code explained
- In simple words: four increasingly forgiving ways to pull nine fields off an invoice, one way to ask a vision model for them, and a set of arithmetic checks that catch many wrong answers whichever way they were produced.
- What happens:
pdf_text()wraps pypdf's extraction, plain or layout.parse_v1()is the first attempt anyone writes: regexes for "Label: value" written while looking at one classic invoice. It works on that layout and nothing else.parse_v2()handles both layouts: the invoice number by its own pattern (INV-dddd-dddddd) wherever it appears, label synonyms (IssuedorInvoice date,Amount dueorTotal due), and flexible whitespace. It expects layout-mode text.ocr_lines()runs RapidOCR, which returns boxes with text, then rebuilds reading order: sort boxes by vertical center, group boxes whose centers are within 60% of a line height into one row, sort each row left to right, and join with four spaces. Without this step, OCR output is a bag of fragments.label()builds a regex for a label that tolerates OCR damage: optional spaces,l/I/1/!/|treated as the same character, stray punctuation after.parse_v3()uses it for every label, accepts a colon as a decimal point, andsnap()s the customer name to the closest known account withdifflib(in production, the list comes from your CRM).money()finds the first amount and treats the last separator as the decimal point, so552,00,144:00, and1,320.00all normalize correctly.best_account()slides each known account name over the page's letters and returns the best fuzzy match above 0.8.parse_v4()is line-based: for each OCR line it takes the letters before the first digit as a label key, fuzzy-matches the key againstLABELS(soAmiountdue,Totaidue, andSubtota!all match), and reads the value after it. The invoice number accepts1NVandlNV, the date falls back to the only ISO date on the page, and seats come fromx N usersor the Qty column. Why v4 exists is the subject of the failure diagnosis below.image_message()builds the OpenAI-compatible message with a text part and a base64 image part.VLM_PROMPTasks for JSON with exactly our field names,nullrather than guesses, and fixed formats.parse_vlm_json()takes the first{...}block (models often wrap JSON in code fences) and coerces types to match the other paths, so every path is scored the same way.vlm_extract()takes anychat-compatible callable, which is how the stand-in and the real model share one code path.check_invoice()runs checks that need no ground truth: all fields present, seats x unit price = subtotal, subtotal + tax = total, and a date format. They cannot prove a value right (a misread that keeps the arithmetic consistent passes) but, as measured below, they catch most wrong documents.main()runs every path on the two layouts and runs the VLM path throughScriptedLLM.
- Comes out: (about 4 seconds with a cold OCR engine)
OCR on scans, and a failure diagnosed
v3 missed the plan on INV-2026-004512 even though "Team plan" is plainly printed on it. Before changing any regex, read what the parser actually received:
"""Read the OCR text before touching the parser: why did v3 miss fields on these scans?
Run: PYTHONPATH=. python examples/m12_ocr_peek.py
"""
from __future__ import annotations
import json
from pathlib import Path
from examples.m12_extraction_eval import cached_ocr
from examples.m12_invoice_pipeline import FIELDS, INVOICES, parse_v3
TRUTH = {t["invoice_number"]: t for t in json.loads(Path("data/attachments/ground_truth.json").read_text())["invoices"]}
for number in ("INV-2026-004512", "INV-2026-004535", "INV-2026-004549"):
text = cached_ocr(INVOICES / f"{number}_scan.png")
got = parse_v3(text)
missed = [f"{f} (want {TRUTH[number][f]!r})" for f in FIELDS if str(got[f]) != str(TRUTH[number][f])]
print(f"--- {number}: v3 missed {', '.join(missed) or 'nothing'}")
print(text.encode("ascii", "backslashreplace").decode()) # show stray non-ASCII OCR characters as escapes
Code explained
- In simple words: print the OCR text that the parser saw next to the fields it missed, so the fix comes from evidence instead of guesses.
- What happens: for three dev invoices, the script loads the cached OCR text, runs
parse_v3(), lists the fields that differ from the answer key, and prints the text with any non-ASCII OCR characters shown as escapes. - Comes out:
BrightlaneTeam plan: OCR dropped the space, so v3's\b(Team|...)word boundary never matches (there is no boundary betweeneandT).552,00: the decimal point was read as a comma, and v3's money pattern only accepts.or:.Arnountdue(mread asrn) andAmiountdue: the label regex allowsl/I/1confusions but not these.Currenicy.: an extra letter.panssi:for "Issued": the label is destroyed entirely.- A middle dot
\xb7appeared from a speck of noise. - These are the classic OCR confusions (
rn/m,l/1/I,./,/:, dropped spaces), and they show why exact-label regexes are the wrong tool for OCR text.parse_v4()changes approach instead of adding more alternatives: compare squashed label keys by similarity, take the last separator as the decimal point, drop the word boundary, and fall back to "the only ISO date on the page" when the date label is unreadable.
v4 now reads all 12 dev invoices perfectly. That is exactly the moment to distrust it: it was written by looking at those 12 invoices.
Scoring against ground truth, with a held-out set
A field-level evaluation compares every extracted field with the answer key and reports accuracy per field, per path, and per document. It is the multimodal version of Module 10's golden set: the "golden" answers come from ground_truth.json. We run it on the 12 dev invoices, which we read while writing the parsers, and on 12 held-out invoices from invoices/heldout/, which no parser version has seen.
"""Field-level extraction eval over the generated invoices (12 dev + 12 held-out), for every path.
Real measurements: the PDF text-layer parsers and the OCR pipeline.
The VLM path runs only when M12_RUN_VLM=1 and a provider key is set.
Run: PYTHONPATH=. python examples/m12_extraction_eval.py
"""
from __future__ import annotations
import hashlib
import json
import math
import os
import time
from pathlib import Path
from examples.m12_invoice_pipeline import (FIELDS, INVOICES, check_invoice, ocr_lines, parse_v1,
parse_v2, parse_v3, parse_v4, pdf_text, vlm_extract)
from examples.m12_vision_chat import VISION_MODELS
from supportdesk.llm import resolve
CACHE = Path("data/attachments/ocr_cache.json")
def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
"""95% Wilson interval for k successes out of n: honest error bars for small samples."""
if n == 0:
return 0.0, 1.0
p = k / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return centre - half, centre + half
def cached_ocr(scan: Path) -> str:
"""OCR is slow on CPU (seconds per page), so keep results keyed by a hash of the file's bytes."""
cache = json.loads(CACHE.read_text()) if CACHE.exists() else {}
key = f"{scan.name}:{hashlib.sha256(scan.read_bytes()).hexdigest()[:16]}"
if key not in cache:
cache[key] = ocr_lines(scan)
CACHE.write_text(json.dumps(cache, indent=1))
return cache[key]
def score(extracted: list[dict], truth: list[dict]) -> dict:
per_field = {f: sum(str(e[f]) == str(t[f]) for e, t in zip(extracted, truth)) for f in FIELDS}
docs_perfect = sum(all(str(e[f]) == str(t[f]) for f in FIELDS) for e, t in zip(extracted, truth))
flagged_wrong = sum(bool(check_invoice(e)) for e, t in zip(extracted, truth)
if any(str(e[f]) != str(t[f]) for f in FIELDS))
return {"per_field": per_field, "docs_perfect": docs_perfect, "flagged_wrong": flagged_wrong,
"false_alarms": sum(bool(check_invoice(e)) for e, t in zip(extracted, truth)
if all(str(e[f]) == str(t[f]) for f in FIELDS))}
def evaluate(truth: list[dict], folder: Path, vlm: bool = False) -> tuple[dict, dict]:
"""Run every path over one split. Returns ({path: extracted}, timings)."""
pdfs = [folder / f"{t['invoice_number']}.pdf" for t in truth]
scans = [p.with_name(p.stem + "_scan.png") for p in pdfs]
started = time.perf_counter()
runs = {"pdf text + v1": [parse_v1(pdf_text(p)) for p in pdfs],
"pdf layout + v2": [parse_v2(pdf_text(p, layout=True)) for p in pdfs]}
pdf_ms = (time.perf_counter() - started) * 1000 / (2 * len(pdfs))
started = time.perf_counter()
ocr_texts = [cached_ocr(s) for s in scans]
ocr_s = (time.perf_counter() - started) / len(scans)
for name, parser in [("scan OCR + v2", parse_v2), ("scan OCR + v3", parse_v3), ("scan OCR + v4", parse_v4)]:
runs[name] = [parser(t) for t in ocr_texts]
if vlm:
model = VISION_MODELS[resolve()[0]] # the default text models cannot see images
runs["scan VLM"] = [vlm_extract(s, model=model) for s in scans]
return runs, {"pdf_ms": pdf_ms, "ocr_s": ocr_s}
def print_table(title: str, runs: dict, truth: list[dict]) -> None:
n_docs, n = len(truth), len(truth) * len(FIELDS)
print(f"{title}: n = {n_docs} invoices x {len(FIELDS)} fields = {n} field values per path")
short = {"invoice_number": "number", "invoice_date": "date", "customer": "cust", "currency": "curr",
"subtotal": "subtot"}
print(f"{'path':16s}" + "".join(f"{short.get(f, f):>7s}" for f in FIELDS) + " fields (95% CI) docs ok wrong docs flagged")
for name, extracted in runs.items():
s = score(extracted, truth)
k = sum(s["per_field"].values())
lo, hi = wilson(k, n)
wrong = n_docs - s["docs_perfect"]
print(f"{name:16s}" + "".join(f"{s['per_field'][f]:>7d}" for f in FIELDS)
+ f" {k / n:6.1%} ({lo:.0%}-{hi:.0%}) {s['docs_perfect']:2d}/{n_docs}"
+ f" {s['flagged_wrong']:2d} of {wrong:2d}" + (f" ({s['false_alarms']} false alarms)" if s["false_alarms"] else ""))
def main() -> None:
truth = json.loads(Path("data/attachments/ground_truth.json").read_text())
vlm = os.environ.get("M12_RUN_VLM") == "1"
dev_runs, timing = evaluate(truth["invoices"], INVOICES, vlm)
print_table("DEV (the invoices we read while writing the parsers)", dev_runs, truth["invoices"])
held_runs, _ = evaluate(truth["invoices_heldout"], INVOICES / "heldout", vlm)
print()
print_table("HELD-OUT (never looked at while writing the parsers)", held_runs, truth["invoices_heldout"])
print(f"\nTime per document: PDF parse {timing['pdf_ms']:.0f} ms; OCR {timing['ocr_s']:.1f} s "
f"({'from cache' if timing['ocr_s'] < 0.5 else 'computed now'})")
misses = [(t["invoice_number"], f, e[f], t[f]) for e, t in zip(held_runs["scan OCR + v4"], truth["invoices_heldout"])
for f in FIELDS if str(e[f]) != str(t[f])]
print("\nHeld-out OCR + v4 errors (invoice, field, got, want):")
for row in misses:
print(" ", row)
if __name__ == "__main__":
main()
Code explained
- In simple words: run every reader over every invoice, compare each of the nine fields with the answer key, and report how often each reader is right, with honest error bars.
- What happens:
wilson()computes a 95% Wilson score interval for k successes out of n, the confidence interval Module 10 recommends for small samples because it behaves well near 0% and 100%.cached_ocr()keys OCR text by file name plus a hash of the bytes, so reruns are fast and a changed file is re-read.score()counts correct values per field, perfect documents, wrong documents thatcheck_invoice()flags, and correct documents it flags anyway (false alarms).evaluate()runs the PDF paths and the three OCR parsers on one split (plus the VLM withM12_RUN_VLM=1).print_table()prints the per-field counts, overall field accuracy with its interval, perfect documents, and how many wrong documents the checks caught.main()scores dev, then held-out, and lists v4's held-out errors.
- Comes out: (about 30 seconds on a cold cache, mostly OCR at about 1 s per page on 2 cores; 0.9 to 1.2 s across our runs. On a rerun the OCR line says "from cache")
- The text layer wins outright when there is one: 108 of 108 fields on both splits, about 2 ms per document, no model. Always try it first.
- v1's 50% is really "100% on one layout, 0% on the other". Averages hide layout-specific failure; per-document and per-layout breakdowns show it.
- Overfitting, measured. OCR + v4 scores 100.0% on dev and 91.7% (95% CI 85% to 96%) on held-out, and perfect documents drop from 12 of 12 to 4 of 12. The dev score measured how well we fit those 12 invoices, not how well v4 reads invoices. Report the held-out number. The v3 to v4 improvement on held-out (77.8% to 91.7%) is probably real: the intervals only touch at 85%, and v4 is better or equal on every field. But with 108 field values per split, differences of a few points are within noise.
- Look at the held-out errors. Every one of them is
None: v4 abstained rather than inventing a value. Read the OCR text for each (cached_ocr()on the held-out scans) and the causes are new confusions:Tearnandpianfor Team and plan,iNvandINvin the invoice number (v4 only allowed the first letter to vary, and only to1,l, or|), the number ending inginstead of9, and the "Amount due" label read upside down asanp yunouiv. The two seat misses are both single-digit quantities (6 and 9) that OCR dropped entirely: the text detector skipped a lone small character. A v5 could fix the case errors with a case-insensitive pattern; no parser can recover a digit that is not in the text. You could derive seats as subtotal divided by unit price, but then the arithmetic check that validates seats is checking itself. Better to leave itNoneand let the check route the invoice onward. And if you do write a v5 from these errors, this held-out set has just become a dev set: generate a fresh one (new seeds inm12_make_assets.py) before you report v5's number. - The checks carry the pipeline. On every path and both splits,
check_invoice()flagged every wrong document (for example 8 of 8 for v4 on held-out) with zero false alarms. That is what makes an imperfect reader safe to deploy: auto-accept documents that pass the checks, send the rest to a human or a stronger reader. It depends on this data, though: a misread that keeps the arithmetic consistent (say, both subtotal and total misread the same way) would pass. Sample some accepted documents for human review too.
A note on honesty: these invoices are synthetic, two layouts, from one generator. Real invoices come from many vendors in many layouts, in several languages (Brightlane's tickets include Spanish, German, Japanese, and Hindi), on crumpled paper. Treat these numbers as a demonstration of the method. Before trusting any pipeline on real attachments, label 50 to 100 real ones (with customer consent and PII handling from Module 11) and rerun this exact script on them.
The VLM path
The third reader sends the scan to a vision model with VLM_PROMPT and parses the JSON. The same eval scores it once you have a key:
export GEMINI_API_KEY=your-key-here
LLM_PROVIDER=gemini M12_RUN_VLM=1 python examples/m12_extraction_eval.py
Code explained
- In simple words: the same exam, with a vision model added as a fourth reader.
- What happens:
evaluate()adds ascan VLMrow by callingvlm_extract()on each of the 24 scans with the provider's vision model fromVISION_MODELS(the default text models on Groq and Ollama cannot see images). Each call costs about 1,120 image tokens plus about 100 prompt tokens and about 150 output tokens on Gemini 3, so the 24 scans cost under 0.10 USD atpricing.py's rates. Mind the free-tier rate limits (Module 13 covers backoff). - Comes out: a
scan VLMrow in both tables, in the same format. We did not run it for this build, so there is no number to show and we will not invent one. When you run it, read the rows the same way: field accuracy with its interval on held-out, perfect documents, and, most important for a VLM, how many wrong documents the checks flag. A VLM tends to fail differently from OCR: instead ofNone, it can return a confident, well-formatted, wrong value (a digit transposed, a plausible total). The arithmetic checks are what catch those.
Here is how a VLM's reply to one scan would flow through the code. Illustrative sample run (not captured in this build; produced for teaching). Your output will differ.
{"invoice_number": "INV-2026-005528", "invoice_date": "2026-04-11", "customer": "Mendez Logistics",
"plan": "Team", "seats": 9, "subtotal": "108.00", "tax": "0.00", "total": "108.00", "currency": "USD"}
Code explained
- In simple words: what a good answer for the held-out invoice whose seat count OCR dropped would look like.
- What happens:
parse_vlm_json()would load it, keep the nine fields, and leave types as they are (already strings and an integer).check_invoice()would confirm 9 x 12.00 = 108.00 and 108.00 + 0.00 = 108.00. - Comes out: in this sample, the model reads the quantity that OCR's text detector skipped, because it looks at the whole region rather than detecting isolated characters. Whether your model does that on your scans is exactly what the eval row tells you.
Vision versus parsing: the decision
| Situation | Use this | Why |
|---|---|---|
| Born-digital PDF from a known system (Brightlane's own invoices) | Text layer plus a layout-aware parser | Measured 108/108 on both splits, 2 ms, free, deterministic |
| Scanned or photographed document, known layouts, high volume | OCR plus a tolerant parser plus arithmetic checks, escalate what fails | Measured 91.7% of fields held-out with every wrong document flagged; about 1 s of CPU per page, no per-call fee |
| Unknown or varied layouts, handwriting, stamps, low volume | A VLM with a JSON schema, plus the same checks | Generalizes across layouts without a parser per template; costs about 1,100 to 2,000 input tokens per page |
| Exact identifiers (invoice numbers, request IDs, amounts) | Parser or OCR first; VLM only as a fallback or a second opinion | Character-exact output; VLM misreads look plausible |
| "What is this document and what does the customer want?" | A VLM (or the text layer plus a text LLM) | Interpretation, not extraction |
| Documents with sensitive data you may not send to a third party | Local parser and OCR, or a local vision model through Ollama | Nothing leaves your machine |
| Anything that triggers money movement (refunds) | Any reader, but checks plus human approval (Module 8, Module 11) | No reader is right 100% of the time |
A hybrid is usually best: the cheap deterministic reader first, checks on its output, and the VLM only for what fails the checks. The lab at the end of this module builds exactly that router.