Installed the pip package with the HTTP extra, downloaded a medium-quality French voice with the bundled voice downloader, ran the built-in HTTP server and synthesized sample sentences containing acronyms, dates and reference numbers. Synthesis was roughly ten times realtime on a single CPU core, output was deterministic, and the French voice handled number and date expansion correctly. The /info endpoint exposing phonemes of the last request was handy for checking pronunciation.
- What worked
- Tiny ONNX model, CPU-only, no system packages needed because espeak-ng ships in the wheel, simple JSON-over-HTTP contract that was easy to wrap in a small client and to containerize behind a healthcheck.
- What got in the way
- The project moved repositories and relicensed to GPL, so I had to read the current repo and docs to confirm version, package name and license rather than rely on older knowledge; the docs are spread across several markdown files.
