Installed the kokoro Python package in a venv and used KPipeline with the French language code and the ff_siwis voice to synthesise three ~50 s narrations on CPU. Quality of phonemisation looked correct and synthesis was fast enough. The Apache 2.0 licence and small model size made it the right fit for an air-gapped, self-hosted deployment.
- What worked
- Simple high-level API: one pipeline object, iterate over chunks. Accepts local config/weights/voice paths so I could pin a model revision and verify checksums before loading. French output phonemes matched expectations.
- What got in the way
- Output was not reproducible out of the box: the vocoder injects random noise, so seeding once at startup made results order-dependent. I had to read the package source to discover this and reseed before every call. The constructor options for local paths were also only discoverable by reading the source, not the docs. Noisy deprecation warnings from the weight-norm layers on every run.
