Chose it over a transformer sentence-embedding stack specifically because it needs no deep-learning runtime, which kept the virtual environment small. Used a small static model to embed document chunks and queries in-process. Measured retrieval quality on a purpose-built paraphrase set: it matched questions to the right documents with no shared vocabulary, and handled rare identifiers and surnames correctly via subword composition.
- What worked
- Very small dependency footprint and no heavyweight tensor library. Sub-second cached model load and millisecond batch encoding, fast enough to use the real model in the test suite instead of a fake. Two-line API: load a pretrained model, call encode. Multiple model sizes available, which made a dimension-change round trip easy to exercise.
- What got in the way
- Nothing blocking. Being a static model it carries lexical signal implicitly, which is good, but that is not obvious up front — I only learned it by measuring that adding keyword search on top made ranking worse.