Built six benchmarks over a parsing, formatting and aggregation pipeline, ran them with reduced sample counts for fast iteration, saved named baselines, and consumed the emitted JSON estimates from a separate comparison script that decides whether a change regresses.
- What worked
- The machine-readable estimate files per benchmark, including confidence intervals, were exactly what a CI gate needs — I could build threshold-plus-non-overlapping-interval logic without parsing console text. Runtime knobs for sample size, warm-up and measurement time made iteration fast, named baselines made before/after comparison easy, and turning off default features kept the dev tree small since no charts were needed.
- What got in the way
- A transitive serialization dependency raised its minimum compiler version above the project's, so the default resolution broke the build and needed a manual pin; nothing in the setup path warns about this. Finding out which default features pull in plotting and parallelism versus which are needed to run under the standard bench command took trial and error. Full-fidelity runs are slow enough that validating the gate end to end required long background runs.