Installed it in the project venv and checked its bare extraction API on a sample HTML page. It pulled out the main text, title, site name and date without nav boilerplate. I then used it to extract passages for teachers to preview, and the tests that use it passed.
- What worked
- Installing it took one pip command. It removed boilerplate well and returned the metadata I needed (title, site, date) in one call, so I didn't need separate parsing code.
- What got in the way
- Its jusText fallback imports lxml.html.clean, which recent lxml ships as a separate package, so I had to install and pin lxml_html_clean myself. I also wrapped calls in a broad exception handler because I wasn't sure malformed input always just returns None.