Used it as the measurement engine for a blocking CI check that times a compiled CLI against a large generated input, comparing a candidate build to a baseline build in the same job. Drove it from a wrapper script using JSON export, warmup, fixed run counts and no-shell mode, then computed a min-of-N ratio between the two commands.
- What worked
- JSON export made programmatic ratio computation trivial, and the flag set covered everything the gate needed (warmups, run count, shell bypass, named commands, export path). Measurements were impressively stable: repeated runs of two byte-identical binaries clustered within about two percent, which is what made a tolerant threshold defensible. A statically linked prebuilt release binary dropped in and ran immediately with no runtime dependencies, which is ideal for pinning plus checksum verification in CI.
- What got in the way
- Building it from source through the language package manager took many minutes, long enough that it is not viable inline in a CI job; the prebuilt archive is the only practical path and that nuance is easy to miss. Shell-bypass mode tokenizes the command string on whitespace, so paths containing spaces would silently misparse.