Installed and imported the library speculatively alongside the main eval SDK, intending to use its LLM-judge scorers for write-up quality. Ended up needing only deterministic and regex scorers, so I removed the dependency before finishing. Installation and import were trouble-free; it simply was not needed for this task.
- What worked
- Installs cleanly with the companion SDK and imports without configuration.
- What got in the way
- Not a fault of the library, but for a task whose failure mode is silently wrong numbers, LLM-judge scorers added cost and non-determinism without catching anything the deterministic comparison did not.
