Skip to content
OPQAI.
Sourced advanced / 💻 Coding Free tools

Boost Search Accuracy with Multi-Vector Embeddings in Sentence Transformers

Job to be done: Implement advanced semantic search and retrieval using multi-vector embeddings with Sentence Transformers

🇳🇬 Ways to use this in Nigeria

Ideas to get you started, adapt to your situation.

  • Student

    For your final year project, build a precise semantic search engine for academic papers, accurately retrieving specific research methodologies or code snippets.

  • 9-5 employee

    As a data scientist, enhance your company's internal knowledge base search to precisely find specific technical specifications or project details across thousands of documents.

  • Entrepreneur

    As a founder building a legal tech platform, implement a search feature that precisely matches user queries to specific clauses in Nigerian legal documents.

What this is, in plain English

This workflow introduces multi-vector embedding models, also known as late-interaction or ColBERT-style models, which are now supported in the Sentence Transformers library (version 6.0 and later). Unlike traditional “dense” embedding models that compress an entire text into a single summary vector, multi-vector models keep a separate vector for each important word or “token” in the text. This preserves more detailed information.

The key difference is how queries are matched against documents. Instead of comparing two single summary vectors, multi-vector models compare each query token’s vector against every document token’s vector. This “late interaction” allows for much more precise matching, especially for specific details, rare terms, or when searching for multiple criteria at once (like “green sofa with wooden legs and rounded cushions”).

While this approach offers stronger retrieval accuracy, it comes at the cost of a larger index (because you store many vectors per document) and potentially slower scoring. This entry is in Concept Mode because implementing these models requires comfort with Python programming, installing libraries, and understanding machine learning concepts like embeddings and indexing. The exact code steps for loading, encoding, and scoring are not provided in the excerpt, but are covered in the full blog post.

What you can use it for

  • Achieve more precise semantic search: Get more accurate results when your search queries contain specific details or multiple requirements.
  • Improve Retrieval Augmented Generation (RAG): Provide Large Language Models (LLMs) with highly relevant information by retrieving documents that match query details more closely.
  • Search visual documents with text: Match text queries directly against page images, without needing to convert the image text into searchable text first (no OCR step).
  • Enhance audio and video retrieval: Apply the same detailed matching approach to find specific moments or content within audio and video files.

Tools you need

  • Sentence Transformers (free): A Python library for creating and using text embeddings and reranker models.
  • Python (free): A popular programming language used to run the Sentence Transformers library.
  • PyLate (free): A Python library/framework whose model checkpoints can be loaded into Sentence Transformers.
  • colpali-engine (free): A Python library for visual document retrieval, whose models can be used with Sentence Transformers.

How it actually works

To use multi-vector embedding models, you will typically follow these general steps within a Python environment:

  1. Install the library: You’ll need to install the Sentence Transformers library using Python’s package manager.

    pip install -U sentence-transformers
  2. Load a multi-vector model: You would then load a pre-trained multi-vector model into your Python script. The excerpt mentions that PyLate and Stanford-NLP ColBERT checkpoints, as well as colpali-engine models, can be loaded. The author doesn’t share the exact code for loading a model; you would typically use a function like SentenceTransformer() with the model’s name or path.

  3. Encode your documents and queries: Convert your text documents and search queries into multi-vector embeddings. This involves processing each text through the loaded model to get a set of token-level vectors. The author doesn’t share the exact code for encoding; this usually involves a method like model.encode().

  4. Score with MaxSim: When a query comes in, you would compare its token vectors against the document’s token vectors using the MaxSim operator to calculate a relevance score. This process is more complex than a simple dot product used for single-vector embeddings. The author doesn’t share the exact code for scoring; this would involve specific functions provided by the library.

  5. Integrate into a search stack: For practical use, these embeddings would be stored in a specialized index (like a vector database) that supports multi-vector search, allowing for efficient retrieval of relevant documents.

The precise code examples and detailed instructions for each of these steps are found in the full blog post and the official Sentence Transformers documentation.

Words you’ll see, explained

  • Embedding: A numerical representation of text (or other data) that captures its meaning, allowing computers to understand and compare it.
  • Vector: A list of numbers that represents an embedding in a mathematical space.
  • Token: A small unit of text, usually a word or part of a word, that a language model processes.
  • Late Interaction: A method where the detailed comparison between a query and a document happens at the scoring stage, after both have been independently processed into token-level embeddings.
  • ColBERT: A specific architecture for multi-vector embeddings that uses late interaction to achieve high retrieval accuracy.
  • MaxSim operator: A mathematical operation used in multi-vector models to calculate the similarity between a query and a document by finding the maximum similarity between query tokens and document tokens.
  • Semantic Search: A type of search that understands the meaning and context of words, rather than just matching keywords, to provide more relevant results.
  • Retrieval Augmented Generation (RAG): An AI technique where a language model retrieves relevant information from a knowledge base before generating a response, making the response more accurate and informed.
  • Checkpoint: A saved state of a trained machine learning model, including its architecture and learned weights, which can be loaded and used.

Original source

This concept is introduced in a blog post by Hugging Face, authored by Tom Aarsen, Antoine Chaffin, and Raphael Sourty. It details the addition of multi-vector (late interaction) embedding models to the Sentence Transformers library with its v6.0 update.

Notes & variations

  • Do you even need this?: For simpler semantic search tasks or when index size and speed are top priorities, traditional dense embedding models (which produce a single vector per document) might be sufficient and easier to implement. Multi-vector models are best for tasks requiring high precision, especially with complex queries or visual data.
  • Free-tier limits: The Sentence Transformers library itself is free to use. However, multi-vector models generate larger indexes and require more computational resources (RAM and CPU/GPU) for both encoding and scoring compared to single-vector models. This means you might need more powerful hardware or cloud resources for large datasets, which could incur costs.
  • Common pitfall: The main trade-off with multi-vector models is the increased storage requirement for the index and potentially slower query times due to the more complex MaxSim scoring. Ensure your infrastructure can handle the larger index size and the computational demands of late interaction scoring before committing to this approach for very large-scale applications.

Keep going

More Coding workflows