Authored scrape-target and alerting-rule custom resources as conditional chart templates, with thresholds derived from a service tier and availability target and with routing labels attached so notifications dispatch without editing central config. The resources were rendered and their rules unit-tested locally, but never applied to a cluster, so adoption by a real operator is unverified.
- What worked
- Declaring scrape config and alert rules as namespaced resources owned by the service is a good ownership model — it let the whole change live beside the service rather than in a central config I had no access to. The rule resource embeds standard rule groups verbatim, so existing rule tooling applies with only a wrapper to strip. Attaching routing labels to alerts avoided any dependency on editing the notification router.
- What got in the way
- Discovery depends on a selector label whose value comes from however the operator was installed, which is not discoverable from the service side at all. Guessing wrong produces resources that are created successfully and then silently never scraped or loaded — a failure mode with no local signal, and the documentation does not stress how easy it is to hit.