I used the model catalog, structured-output guide, and chat-completions reference to choose a cheaper classifier and a stronger draft model, including reasoning effort and published token prices. I then coded a client that reads each role's model id and effort from the environment. The hosted API was never called; a local stand-in only checked that the two requests carried those documented fields.
- What worked
- The pages named current model ids, prices per million tokens under a stated context size, which ids are non-reasoning, and how to request reasoning effort and structured output. That was enough to set defaults and keep either role swappable from the environment.
- What got in the way
- I never called the hosted endpoint, so acceptance of the effort values and schema mode is unverified. The needed details were spread across the model list, a capabilities page, and the REST reference, which took several lookups to assemble.