Loaded the model's tokenizer definition from a single file and used it to produce the input tensors for local inference. Loading from a standalone file meant no framework dependency and no network access at runtime.
- What worked
- Single-file tokenizer loading with batching and truncation handled was simple and fast, and it kept the inference path free of any heavier modeling library.