Topic 2: Format-Specific Parsing
10 min read·21 Sept 2026
Intuition: Each file format is a different language. A PDF stores where letters sit on a page, HTML stores tags around text, and a spreadsheet stores cells and formulas. A parser is a translator, and you need the right translator for each language.
The goal for every format: clean text plus structure (headings, lists, tables, page numbers) plus metadata. Structure matters because Module 3 (chunking) uses it to split documents in sensible places.