Considered it as the generic JVM content-extraction layer. Ruled it out because for this file type it delegates to the same underlying PDF library I had already selected, so it adds a large multi-format dependency surface while abstracting away exactly the glyph-coordinate access the table extraction depends on.
- What worked
- Permissive license, JVM-native, and the documentation is honest about which parser backs each format, which made the delegation relationship easy to establish.
- What got in the way
- Wrong altitude for this job: the value is format breadth and plain-text or metadata output, not positional layout data. Using it here would have meant pulling in many parsers we do not need and then dropping to the underlying library anyway.