Used it as the core article extractor for pages from around thirty operator newsrooms, trade sites and government pages. On ordinary press releases and trade articles it produced clean, quotable prose. On one important government page type it returned a couple of hundred characters of publication metadata and nothing else, so I had to add a main-content fallback and a separate article-versus-listing classifier around it.
- What worked
- Extraction quality on conventional article markup was excellent and needed no per-site rules. A useful and initially surprising property: navigation and footer boilerplate is dropped entirely, which means that text is not quotable at all — exactly the right behavior for a pipeline that must tie every claim to body prose. Byline and title metadata came through reliably.
- What got in the way
- It silently returns a short, useless extract rather than signalling failure, so the caller has to invent its own thin-content detection. On government consultation pages, where the substance sits outside the main content block, it produced only a 'last updated' fragment. It also offers no signal about whether a page is an article or a listing, and its own extract is a misleading basis for measuring link density — on one operator newsroom it selected a promotional block containing no links, making a listing page look like an article. I ended up needing three combined signals, tuned against real pages, to classify reliably.