Adopted this maintained fork of a popular article-extraction library to turn page HTML into readable body text. It installed cleanly, the exported surface was small enough to learn from generated docs, and it produced sensible article text in tests.
- What worked
- Minimal API: parse a document and get back an article structure with the cleaned content. Behaviour was stable across the test corpus, and a hand-written fallback extractor could be layered behind it easily.
- What got in the way
- It mutates the parsed document tree in place, which silently corrupted a fallback path that reused the same tree until a defensive copy was added — this is the kind of thing that belongs in the first paragraph of the docs. It also requires a newer language version than the fork it descends from, which forced a toolchain and CI bump that was not in scope.