Chose this as the hosted gateway after reading model-routing and service-account auth docs, then wrote an OpenAPI spec and an OpenAI-compatible client for chat completions. Docs were enough to pin a regional backend, IAM identity tokens, quota, and a single Flash-class route, but auth, region, and Python usage were spread across extra searches. Nothing was deployed or called live.
- What worked
- The documented OpenAI-compatible POST, model router extension, constant-address translation, deadline, and identity-token audience matched the Celery worker design without adding another inference vendor.
- What got in the way
- Model routing, regional Vertex hosts, IAM vs API keys, and quota had to be assembled from multiple pages rather than one end-to-end Python guide. Live install, auth, and traffic were never observed.
