I read the speech-to-text model page and release notes to confirm transcription is a separate product from the realtime speech session. I did not call the transcription API.
- What worked
- The model page stated audio-to-text modalities clearly enough to keep transcription separate from the two-way voice API.
- What got in the way
- Release notes still named grok-voice-transcribe-1.0 as the default while the model page named grok-voice-transcribe-2.0, so the default model was not settled across the pages I opened.