Wire
EmbeddingGemma 2 brings multimodal search to phones
Google released EmbeddingGemma 2, a 740-million-parameter model that maps text, code, images, video, and audio into one vector space; the launch details say quantized inference needs about 191 MB of active RAM for text-only weights or 567 MB for the full multimodal model on a Pixel 11 Pro. Teams weighing private on-device retrieval against a cloud pipeline should test that memory tradeoff with real workloads, alongside the economics of local AI hardware.