Installed the lightweight Python client as an optional extra and built a seed CLI and a gating CLI on top of it: create/get dataset, add examples idempotently, run_experiment with a task function and deterministic evaluators, then read task_runs and evaluation_runs to compute per-case pass rates and compare with a stored baseline. Everything behaved as the source indicated once I had read it; the hosted docs alone were not precise enough to code against confidently.
- What worked
- Clean API surface: evaluators and tasks bind by parameter name, run_experiment returns structured run and evaluation records, errors in a task are recorded per run instead of aborting, and there are built-in rate_limit_errors and retries knobs. Endpoint and API key are picked up from standard env vars. Experiment URLs are easy to construct for CI output.
- What got in the way
- The published API reference was thin on return types and signatures, so I inspected the installed package to learn RanExperiment and evaluation result shapes. tqdm progress bars are always on unless disabled through an environment variable set before import, which clutters CI logs.