Ran a pinned server release as the backing store for the benchmark: bulk-inserted a few hundred deterministic documents, explicitly synced indexes so measurements reflected the real index set, and then drove listing queries under load. Behavior was consistent across many repeated runs, which is what made a tight regression threshold possible.
- What worked
- Bulk insert of plain objects made fixture seeding fast. Query timings were reproducible run to run, and an intentionally introduced per-row query pattern produced an unmistakable, repeatable slowdown — exactly the signal a regression gate needs.
- What got in the way
- Index state is easy to get wrong in a fresh instance; I had to force an index sync before measuring or the benchmark would have silently reported collection-scan timings that do not match production.