Learn · Intermediate
Matryoshka representations: useful embeddings at several sizes
Matryoshka representation learning trains an embedding so that selected prefixes of its coordinates remain useful representations in their own right. A search system can keep a short vector when storage and speed matter, or a longer one when accuracy matters more, using the same trained encoder. The method changes how information is organized in a representation; it does not shrink the encoder’s weights.
An ordinary embedding places an item in a space where similar items should be nearby. If a model returns 768 numbers for each document, the search system stores those coordinates and compares a query against them. At millions of documents, vector storage and comparison work become substantial. The question is whether every use of the representation needs all 768 numbers.
The answer is not automatically yes or no. In an ordinary vector, important information may be spread across all coordinates. Throwing away the final two-thirds can destroy distinctions the model learned. It is like deleting the last two-thirds of every sentence and hoping the beginning contains every important qualification. Mechanical truncation is possible; useful truncation is a training property that must be established.
Aditya Kusupati and colleagues introduced Matryoshka Representation Learning to create nested useful representations. The name refers to nesting dolls: a smaller doll is contained inside a larger one. Here the first small set of coordinates is contained inside a longer vector, which is itself contained inside the full vector. Each supported size is trained to perform the relevant task.
The mechanism is straightforward. During training, the encoder produces its full representation. The training procedure also takes several chosen prefixes, such as the first 64, 128, and 256 coordinates. A task loss is applied at each selected size, and those losses contribute to the same parameter update. The encoder is rewarded for making the beginning useful while allowing later coordinates to add detail.
Imagine a map that starts with a country, then adds a city, then a street and building. Someone routing international shipments may need only the broad information; a courier needs the fine detail. That is an analogy for progressive usefulness, not a claim that particular embedding coordinates literally encode countries or human-readable concepts. The nested vectors remain learned numerical features.
This technique can be used with different training objectives. In classification, shorter representations must still support the class decision. In retrieval, they must still place matching items near their queries and separate relevant alternatives. Contrastive learning supplies one way to train those relationships. Matryoshka learning adds the requirement that several prefixes work, rather than changing the fundamental meaning of similarity.
A concrete example comes from Google’s EmbeddingGemma 2 release. Its full output has 768 dimensions. Google reports close-to-lossless quality down to 256 dimensions, while the model card warns that 128 dimensions substantially degrades multimodal quality. The result belongs to that model and evaluation; 256 is not a universal safe size for every task or dataset.
Keeping 256 rather than 768 coordinates uses one-third as many stored coordinate values. That is an arithmetic reduction in the vector payload at the same numeric precision, not a promise that the entire index is one-third as large. An index also stores document identifiers, metadata, graph links, and other overhead. Approximate nearest-neighbor search may introduce its own storage and accuracy tradeoffs.
Normalization matters too. If the full vector had unit length, its truncated prefix usually does not. Google instructs developers to normalize truncated vectors again before cosine-similarity comparisons. The query and stored items must use compatible representations and dimensions. A short query cannot simply be compared to a longer corpus vector without an explicitly compatible operation.
The EmbeddingGemma developer guide makes the deployment distinction concrete: modular encoders determine which input types can be represented, while vector truncation determines how much of the output is retained. Leaving out an unused media encoder and shortening a search vector are different resource decisions. Neither should be described by a single ambiguous model-size number.
A system can also retrieve candidates with a short representation and rerank a smaller set with a longer one. Whether that saves time depends on candidate recall and the implementation. If the short search discards the right document, the long reranker cannot recover it. Developers should measure retrieval quality at each intended size, especially for rare categories, languages, or cross-media queries.
The honest limit is that nesting preserves usefulness approximately, under a training objective and measured distribution. It cannot guarantee every fine distinction survives every compression. Matryoshka learning offers a principled storage-quality choice; quantization changes numerical precision instead. Understanding that separation helps readers ask the right deployment question: which representation size preserves the information this particular search task needs?
Key questions
Can any embedding be shortened by keeping its first coordinates?
Does shortening an embedding reduce the model download?
Why normalize a vector again after truncating it?
Cite this
APA
Ground Truth. (2026, October 7). Matryoshka representations: useful embeddings at several sizes. Ground Truth. https://groundtruth.day/learn/matryoshka-representation-learning.html
BibTeX
@misc{groundtruth:matryoshka-representation-learning,
title = {Matryoshka representations: useful embeddings at several sizes},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/learn/matryoshka-representation-learning.html}
}