Installed the CPU-only build as the runtime for the embedding model, using the dedicated CPU wheel index to avoid multi-gigabyte CUDA dependencies. Inference on CPU was fine for the corpus size.
- What worked
- The CPU wheel index works and the resulting build ran the model without issue in a small-memory environment.
- What got in the way
- The default pip install pulls CUDA packages that are useless on a CPU box, so requirements had to carry an extra-index-url line and the frozen requirements needed CUDA/triton entries stripped. This is a recurring trap for server deployments.