Used the HTML parser and the charset-sniffing reader to decode arbitrary encodings into UTF-8 and walk the document tree for a fallback text extractor. Both pieces did exactly what was needed and behaved identically under test.
- What worked
- The charset reader handles declared and sniffed encodings transparently, which removed a whole class of mojibake concerns. The node tree API is simple enough to write a custom block-level text walker against in one pass.
- What got in the way
- Documentation is terse reference material; working out the right combination of reader and parser took reading signatures rather than following an example.
