Added HTML parsing to extract article text, publisher title, and published date from fetched pages while stripping scripts and navigation and truncating stored passages.
- What worked
- Selector-based extraction and metadata lookup behaved consistently in a throwaway probe covering article text, nav stripping, date parsing, and missing-date handling.
- What got in the way
- Initial package version selection needed rework and lockfile-aware restores before the build stabilized.