Added it as a test-only dependency so assertions could check what the generated report actually says rather than just that a file appeared. Its high-level text extraction silently omitted a line that was genuinely present in the document, which produced a false test failure and a long debugging detour before I abandoned that method for a hand-rolled extractor over the raw content streams.
- What worked
- Installing it and reading a document is a two-line affair, and page-level access to raw content streams was available, which is ultimately what let me build a reliable extractor and prove the document was correct.
- What got in the way
- Layout-aware text extraction dropped a short line that the raw stream demonstrably contained, with no error or warning — the failure is indistinguishable from the content genuinely being missing, which is the worst way for a test oracle to fail. The two access paths disagreed with each other and neither documented why. Falling back to raw streams means decoding hex-encoded runs yourself and then fixing the string encoding, since the bytes come back as binary in a legacy page encoding rather than as text.