The job is scripted to start llama-server on loopback only, leave GPU offload off, poll until it is healthy, and request a completion without an HTTP proxy. Weights are expected to already be on the runner. I used the current GPU-layer flag after a recent rename. The binary and weights were not on this machine, so the server was never started.
- What worked
- A local completion server with an explicit bind address and a switch to disable GPU offload fits a CPU-only runner that must not call an external model API.
- What got in the way
- Startup, health checks, flag compatibility, and output quality were not observed because the server binary and weights were absent. The job assumes both are installed beforehand.