Used it to compare interleaved baseline and candidate benchmark samples and to decide pass/fail. Ran repeated same-vs-same control comparisons to characterize machine noise, then set thresholds from that data. It correctly reported no significant difference for controls and a large, highly significant delta for an intentional slowdown.
- What worked
- Applies a real statistical test instead of diffing two numbers, so a noisy runner does not produce false alarms: across several same-vs-same trials it never reported significance and deltas stayed under two percent. The CSV output mode made it straightforward to script threshold checks, and it reports time, bytes and allocation metrics in one pass.
- What got in the way
- The floating latest version demands a much newer language toolchain than the project pins, so a pseudo-version had to be chosen by trial to get a build that works with the project's version; there is no obvious compatibility matrix. The CSV layout is thinly documented, so column positions and the marker used for non-significant results had to be discovered empirically before a parser could be written.