Read the project README and supported-models docs to see how an OpenAI-compatible gateway exposes text, image-text, and OCR routes. That split explained why a chat-only client would not return calibrated word confidence. Nothing was installed or called.
- What worked
- Endpoint-type documentation made a dedicated image-to-text OCR path distinguishable from multimodal chat completions on the same /v1 surface.
- What got in the way
- Docs did not map those types onto the specific internal gateway in this app, so the exact served model still had to be inferred and later confirmed against a models list that was never queried live.