Adrien Grondin, Founder of Locally AI, demonstrated how to run Google's Gemma 4 large language model on an iPhone using Apple's MLX framework. This development signifies a significant step towards enabling powerful AI capabilities directly on mobile devices, bypassing the need for cloud connectivity for certain tasks.
Adrien Grondin and Locally AI
Adrien Grondin, founder of Locally AI, presented the demonstration. Locally AI is a startup focused on bringing AI models to on-device applications, making AI more accessible and efficient for end-users. Grondin's work highlights the growing trend of democratizing AI by enabling its deployment on consumer hardware.
Running Gemma 4 on iPhone with MLX
The core of the demonstration revolved around running Gemma 4, a family of open models developed by Google DeepMind, on an iPhone. Grondin showcased how to use the MLX framework, developed by Apple, to optimize these models for Apple Silicon chips found in iPhones and Macs. MLX is designed to facilitate efficient on-device machine learning tasks, including natural language processing and image generation.
Grondin explained that while the full-sized Gemma 4 model can be resource-intensive, quantized versions are available and perform exceptionally well on mobile hardware. He specifically mentioned the utility of 4-bit, 6-bit, and 8-bit quantized models, which offer a balance between performance and accuracy, making them suitable for on-device applications.
Performance and Quantization
The presentation highlighted the speed and efficiency achieved by running Gemma 4 on an iPhone. Grondin stated that the model can process approximately 40 tokens per second. He noted that this performance is achievable even on older iPhone models, indicating the optimization capabilities of MLX.
