I added PDFBox 3.0.8 and used it in-process to read each page's text layer, group characters into rows, and post only when minor-unit line totals matched the payment. A large multi-page fixture, European amount formats, and parenthesized credit amounts passed. Row assembly was custom code, commons-logging had to be excluded for the app's logging bridge, and non-embedded standard fonts needed system fonts plus a font-cache setting.
- What worked
- The 3.0.8 API loaded documents and exposed ordered text positions with no network call, under the Apache License 2.0. Once fonts were present, extraction was stable enough for exact integer sums and for rejecting textless or unbalanced advices before any posting.
- What got in the way
- PDFTextStripper does not return tables, so baselines had to be grouped in application code, and the writeString override is easy to get wrong. Tests logged a LiberationSans fallback because standard fonts were not embedded; without those fonts the same text can extract as empty. The logging bridge also kept that warning visible.