Installed this as the one new runtime dependency to pull structured-data blocks, meta tags and titles out of arbitrary publisher HTML. Wrote a small probe first to pin down the two behaviours I could not afford to get wrong: whether attribute values come back already entity-decoded, and whether markup inside HTML comments is correctly excluded. Both behaved exactly as needed, and the parsed output then held up against several real-world pages fetched live.
- What worked
- Tiny install footprint for what it does, a minimal and obvious API surface, and correct handling of the messy cases that make hand-rolled regex parsing a trap - commented-out tags, quoting variants and attribute ordering. Attribute access returns decoded values, which avoided a double-decode bug that would have been permanently baked into stored text.
- What got in the way
- The decoding behaviour of attribute access is not something I was willing to take from the README alone, so I had to verify it empirically. Clearer documentation of entity handling would have saved that step.