Loaded a pinned revision of an open-weight BERT-style embedding model in-process on CPU, produced normalized passage and query embeddings with the model's recommended query prefix, and reused its tokenizer for offset-aware chunking. Semantic search tests passed against the real model.
- What worked
- Loading by model id plus revision, normalizing embeddings, and encoding batches took very little code. Running on CPU within a 3 GB memory budget was fine for a base-size model.
- What got in the way
- The dependency footprint is heavy: it pulls torch, transformers and scikit-learn, and first run downloads roughly 440 MB of weights. Progress-bar and deprecation noise cluttered test output and had to be filtered.