Chose this library to pull per-page text and a page count so that model-reported page references could be checked against the real text layer. It compiles and integrates cleanly, but I never ran it against a real PDF in this task, so its behavior on actual documents is unverified.
- What worked
- Small, clean API with a proper semantic version tag — a page count and per-page content accessor was all I needed, and wiring it in took one pass after reading the exported signatures. No transitive dependencies, so it added nothing to the dependency graph.
- What got in the way
- Text extraction is primitive: content comes back as positioned glyph records rather than readable strings, so reassembling anything resembling lines is on the caller. It is also known to panic on malformed input rather than returning an error, so every call had to be wrapped in a recover — a library that parses untrusted files should return errors. Documentation is essentially the exported signatures; there is no guidance on reconstructing reading order.