While designing how phone-browser audio should start, I opened the text-to-speech docs several times looking for mobile Safari audio-session rules, including user-gesture and suspended-audio behavior. I did not call the text-to-speech API.
- What worked
- The text-to-speech pages were separate from the speech-to-speech and speech-to-text pages, which helped keep the products distinct.
- What got in the way
- I reopened the same pages and ran several site searches before I could tell whether Safari suspension rules were documented there. Those pages did not become the integration I shipped.