Installed the on-device LLM web package and read dialog, interrupt, and worker types so job questions could be answered from a local payload with conversation state preserved. A console-issued model file is required; the LLM was never loaded or queried.
- What worked
- Dialog and interrupt APIs in the typings matched the need to keep place in a conversation and cancel a stale reply after barge-in. Public writeups made the local-browser fit clear.
- What got in the way
- Weights are distributed as a separate model from the console, not the npm package. Function calling and multi-turn quality were not observed at runtime.
