模型精选73°

Google发布EmbeddingGemma 2模型

精选理由

Google新发布的EmbeddingGemma 2能让手机本地处理多模态数据,无需上传服务器,还特别优化了代码搜索能力。

Google推出EmbeddingGemma 2模型,采用Apache 2.0许可证。该模型拥有740M参数,基于Gemma 4架构,支持文本、代码、图像、音频和视频的多模态处理。在Pixel 11 Pro上,量化文本权重占用191MB RAM,完整多模态模型占用567MB RAM。代码搜索性能显著提升,MTEB Code分数从68.76升至78.68。

原文 · rohanpaul_ai

Google dropped EmbeddingGemma 2 for on-device multimodal AI, under an Apache 2.0 license

> puts text, code, images, audio and video into 1 searchable space on phones. gives phones a missing piece: a way to understand and search your own stuff without sending it to a server.

> 740M parameters, uses the Gemma 4 architecture.

> Its parts are modular, so a text-only app needs just 270M parameters, while a 170M vision encoder and a 300M audio encoder load only when needed.

> On a Pixel 11 Pro, quantized text weights take about 191MB of active RAM, and the full multimodal model takes about 567MB.

> The context window grows 4x to 8K tokens, enough for roughly 5.5 minutes of audio, 29 images or 58 video frames in 1 input.

> Code search gained most, with the MTEB Code score rising from 68.76 to 78.68 while multilingual text scores held level.

> Google also claims top sub-1B results on audio and vision benchmarks and wins over some specialist models twice its size,