
A lightweight open multimodal model family for local apps, with text, image, video understanding, long context, and multilingual support.
Best for: Developers who want an open Google model that can run on a workstation, laptop, or supported mobile and cloud setup.
Build local or edge AI experiences with an open multimodal model that understands text, images, and audio.
Meta open-weight multimodal models for text, image understanding, and long-context AI applications.