Added robots.txt parsing with a per-host cache in front of the page fetcher, so a disallowed URL is refused before any request and the refusal is recorded as a reason rather than silently dropped.
- What worked
- Zero transitive dependencies and a two-call API: parse, then ask whether a user agent may fetch a path. Compatible with the project's older language version without pinning gymnastics. Behaved correctly in unit tests against served and missing robots files.
- What got in the way
- Documentation is minimal — essentially the package comments — so semantics for edge cases like unreachable or malformed robots files had to be confirmed by reading the source and writing tests.