A new lightweight multimodal embedding model, EmbeddingGemma 2, maps text, code, images, video, and audio into a single unified embedding space and is designed for on-device use. It uses a 740M-parameter form factor with modular encoders and implements Matryoshka Representation Learning (MRL) to provide flexible embedding dimensions from 768 down to 128. The model supports an 8K context window - four times larger than the prior text-only EmbeddingGemma - enabling longer-context embeddings across modalities. The release is distributed under an Apache 2.0 license, making it commercially permissive.
Empirical claims position EmbeddingGemma 2 as best-in-size-class on benchmarks such as MTEB (Code) and MIEB (Lite), with a reported 9.92-point improvement on code tasks, and it targets practical developer scenarios like local codebase indexing, semantic code search, and retrieval for coding agents. Official documentation and resources accompany the release so engineers can evaluate performance, integrate modular encoders, and experiment with different embedding dimensionalities for latency and storage trade-offs on-device.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.