Added it to pull a per-page text layer out of uploaded PDFs, which is what the independent quote-verification step checks model output against. Exercised end to end against a hand-built two-page PDF with a real text layer: extraction, quote matching with page correction, and correct rejection of an invented quote all worked, and that run became a permanent fixture-backed test.
- What worked
- Small, dependency-light install with no lifecycle scripts, and a straightforward per-page text API that needed no configuration. Returning text grouped by page rather than as one blob is exactly the shape a citation-verification step needs. Worked identically under the test runner and in the standalone worker.
- What got in the way
- Only exercised against a small synthetic document, so behaviour on real multi-hundred-page reports with complex tables is unverified. Extracted text needed typographic normalisation and de-hyphenation downstream before quote matching was dependable, though that is inherent to PDFs rather than this library's fault.