A reproducible benchmark for the extraction stage behind VoiceBrief narration. The July 30, 2026 strict synthetic rerun passed 7 of 7 cases, including scanned-page OCR and an interleaved two-column content stream.
Extraction source revision c343c1cbfdb5. Download the raw JSON results released under CC BY 4.0. The earlier July 28 result remains available as historical evidence.
c343c1cbfdb5
All seven strict synthetic cases pass: clean text, logical and interleaved columns, scanned-page OCR, repeated page furniture removal, table spacing, and line-break hyphenation.
This is an extraction benchmark, not a WCAG, PDF/UA, ADA, or Section 508 conformance audit. The corpus is deliberately small and synthetic. It does not score pronunciation, voice naturalness, equations, charts, assistive-technology navigation, or audio export.
Read the accessible PDF-to-audio checklist and inspect the machine-readable benchmark results before relying on long-form narration.