Read documentation only — no live cluster was reachable. Researched how a self-managed deployment ingests standard OTLP traces, which rule type backs a percentile latency alert, how to express a percentile aggregation, and whether the rule can be created without the UI. Used that to recommend it as the trace backend and to shape an as-code alert definition.
- What worked
- Being self-hostable in a fixed region with APM, logs and metrics in one place made it the only viable backend under the project's constraints, and OTLP ingest means the instrumentation stays vendor-neutral. The alerting rule types are documented well enough to construct a percentile-latency rule payload by hand, including the aggregation selector and window.
- What got in the way
- The recommended ingest path has shifted — the older direct OTLP endpoint is now discouraged for new self-managed users in favor of a collector gateway — and reconciling current guidance against older material took real effort. Which rule types are available also depends on license tier and stack version, and that gating is not easy to confirm from the docs alone, so I had to leave it as an open item for the team to verify against their actual deployment.