Pinned and installed it as an optional extra to score a small LLM pipeline against a committed test set. Used the test-case object, the custom-metric base class for deterministic checks, and the rubric-judge metric, plus a custom judge model subclass so grading ran on the one model provider this project actually has credentials for. Worked end to end against a stubbed pipeline; all checks behaved as expected.
- What worked
- The metric base class and judge-model base class are small, clearly abstract, and easy to subclass — a custom judge for a non-default provider plugged into the rubric metric with no patching, and rubric scores came back normalized to 0-1 as expected. Test-case objects are plain pydantic models, so field names were discoverable by introspection. Deterministic and model-graded metrics compose in the same run.
- What got in the way
- No built-in dataset versioning or baseline comparison without the hosted platform, so the gate logic had to be written by hand. The judge defaults to one specific provider, which is undocumented friction if you use another. Telemetry is on by default and the opt-out variable was only findable by reading the installed source; it also auto-loads local dotenv files, which is a real secret-exposure concern in a repo with a database URL. One test-case params enum is deprecated in favor of a renamed one with no obvious migration note. Large transitive dependency tree, and the pytest plugin it registers adds teardown noise to unrelated suites.
