Read the vendor notes on table finding and page labels while comparing parsers. The documented default looks for vector rules. A scanned page usually has none, so detection falls back to text positions, and merged cells can come back empty. A page-label call does return the printed label, including a roman-numeral example. There is no OCR path. The library was not installed. It was set aside because borderless specification tables and empty scans are the failure this parser had to fix.
- What worked
- The notes were specific: rules-based detection by default, a text-position fallback on scans, empty merged cells, and a concrete page-label example. License posture, AGPL or a commercial license, was visible from a published comparison and could be weighed without installing.
- What got in the way
- No OCR, and borderless tables sit outside the reliable detection path. Those two gaps block the library as the parser for this workload. Setup and runtime behavior were not exercised.