Omar Sanseviero, Developer Experience Lead at Google DeepMind, recently presented an overview of the Gemma 4 family of open models at the AI Engineer Europe conference. The presentation highlighted the models' capabilities, performance, and the broader impact of open-source AI development. Sanseviero emphasized the rapid progress and broad adoption of Gemma since its initial release, showcasing its versatility across various applications and devices.
Omar Sanseviero's Role at Google DeepMind
As the Developer Experience Lead at Google DeepMind, Omar Sanseviero plays a crucial role in bridging the gap between cutting-edge AI research and practical application. His work focuses on ensuring that developers can easily access, understand, and utilize DeepMind's advanced AI models. Sanseviero's presentation underscores his team's commitment to fostering a thriving AI developer community by providing accessible tools and comprehensive resources.
Introducing the Gemma 4 Family of Models
Google DeepMind recently released Gemma 4, a family of open models designed to be both powerful and accessible. Sanseviero noted that the models were launched just days before his presentation, creating significant excitement. The Gemma family includes several versions, ranging from 1 billion to 27 billion parameters, each offering different trade-offs between performance, size, and computational requirements. These models are designed to be run on various infrastructures, from cloud servers to personal devices, making advanced AI capabilities more widely available.
Gemma Model Sizes and Capabilities
The Gemma 4 models are available in several sizes, including 2B, 4B, 26B A4B, and 31B parameters. Sanseviero presented a table detailing these models, their effective parameters in VRAM, GPU consumption at 8-bits, and their intended use cases. The 2B and 4B models are described as "tiny" and are suitable for edge and on-device applications, capable of running on Android, iOS, Raspberry Pi, and Jetson Nano. The 26B A4B model is noted for its very fast inference, while the 31B model offers maximum quality and fine-tuning capabilities, fitting on a single consumer GPU.
Sanseviero highlighted the efficiency of the 2B and 4B models, stating, "you can run in your own infrastructure, your own devices." He elaborated that these smaller models can perform tasks like multimodal reasoning and on-device inference, demonstrating their practical utility. The larger models, like the 31B, are positioned for more demanding tasks requiring higher intelligence and reasoning capabilities.
