Installed the library and used its feature-extraction pipeline to compare two small sentence-embedding models on sample notes, then made it the in-process embedding engine of a Node/Next.js app. Both models ranked the right note first for every test phrasing. The chosen quantized model ran at roughly 70 ms per query.
- What worked
- Simple pipeline API with pooling and normalization options, quantized model support, and no paid service or API key. It ran the same way in a standalone script, in tsx scripts and inside the Next.js server.
- What got in the way
- The model is downloaded on first use, so the server needs network access and a writable cache directory, and the first search is slow. It also had to be marked as a server external package in the Next.js config.