Installed camelot-py 2.0 and used lattice extraction with a hybrid fallback to pull ruled and borderless tables from in-memory PDFs, then mapped accuracy, page, and column edges into the existing reader types. Pip install was enough; Ghostscript was not required. Header copy helpers and composite confidence did not match this corpus, so adapter code had to compensate.
- What worked
- Lattice mode recovered ruled grids, including column x-edges used to stitch page continuations. BytesIO input worked. Hybrid flavor extracted a borderless text table after lattice returned nothing. Inspecting Table.accuracy, data, and parsing_report made mapping to the existing schema straightforward once the objects were in hand.
- What got in the way
- copy_text and copy_spanning_text did not fill year header cells that were empty rather than truly spanning, so years had to be forward-filled in adapter code. The install-packages documentation page returned 404. Composite confidence mixed in whitespace and scored a valid sparse table below a 0.90 floor, so raw accuracy was used instead.