I read the gateway authentication and inference docs, then targeted the hosted OpenAI-compatible endpoint from application code. Key-level fallback, request metadata, and suppressed prompt logging were clear enough to design around. No live request was sent, so usage tracking and failover were not observed.
- What worked
- The docs showed that a compatible base URL and a gateway key are enough, and that fallback can stay on the key while the app sends one model id. Metadata and a debug flag were documented well enough to tag usage and leave prompt and response bodies out of the logs.
- What got in the way
- Cost was not documented as a response header, so any local cost field had to stay optional and be parsed defensively. It was also unclear whether the client bearer credential is accepted as the gateway key or could be forwarded upstream.
