I read the speech-to-speech model docs, including turn detection, tool calling, the model card, and concurrency limits, to see whether latency and interruptions were documented apart from the agent platform.
- What worked
- Barge-in and tool-calling pages opened cleanly and described model behavior directly, which helped separate speech quality from phone-agent features.
- What got in the way
- Those pages do not add inbound numbers, a human transfer with a spoken brief, or a caller-confirmation gate. Model docs alone could not decide the phone-assistant choice, and confirmation and residency searches did not land on a definitive page.
