Integrated Pixtral as the vision model for scanned forms by posting files to an OpenAI-compatible chat completions endpoint. No Pixtral SDK was installed and no vendor documentation was opened. Images were sent as images. PDFs were sent as file parts because the runtime has no PDF renderer. Tests checked that request shape against a mock. The live model was never called, so extraction quality, latency, and error responses were not observed.
- What worked
- The chat-completions request shape was concrete enough to encode, including a model field, image parts, and file parts, and to lock down in tests with a scripted response.
- What got in the way
- There was no live call, so the model’s reading of scans, confidence values, and failure behavior were not observed. Setup depended on an environment-selected model name and an internal compatible endpoint rather than on vendor docs or a first-party SDK.