Researched on-device wake word, streaming speech recognition, text to speech, voice activity detection and intent options for up to an hour without connectivity. Documentation made clear the engines run on device with only license check needing network, unlike cloud speech pipelines, so it was selected and a local engine boundary was stubbed for later SDK plug-in.
- What worked
- Documentation clearly separated on-device capability from network-dependent steps and mapped engines to needed behaviors such as cached-data answers, confirmation before writes, barge-in handling and conversation resume.
- What got in the way
- Live SDK was never installed, initialized or run against audio in the record, so real-world accuracy, latency and license behavior remain unverified.
