Installed the local model runner, started its server, retrieved review models, and called its local API for diff review. Setup needed extra runtime tuning on constrained hardware and inference was very slow, but all calls stayed local.
- What worked
- Local-only execution and simple model retrieval supported the privacy requirement well.
- What got in the way
- Inference on limited CPU and memory was impractically slow for a full review and needed temporary resource workarounds.