Sakana AI Tests Gemma 4 for Orchestration

Sakana AI validates Gemma 4 for Fugu orchestration, proving base model independence and enabling greater customer choice and sovereignty.

Diagram showing Sakana Fugu's two-layer architecture with conductor and model pool
Sakana
Visual TL;DR
Sakana Fugu ArchitectureCore
unified platform for multi-agent orchestration, intelligently determines processing path for requests
From the article 3 mentionsSakana AI has verified that its Fugu orchestration system can achieve comparable performance using Google's Gemma 4 as the foundation for its conductor model.
Two LayersContext
conductor model dictates collaboration, model pool consists of actual processing units
Gemma 4 VerificationCore
Sakana AI validates Gemma 4 for Fugu orchestration, proving base model independence
From the article 4 mentionsTo rectify this, Sakana AI trained a conductor model using Gemma 4, an open model from a different lineage than their previous Qwen-based conductor.
Comparable PerformanceOutcome
From the article 4 mentionsSakana AI has verified that its Fugu orchestration system can achieve comparable performance using Google's Gemma 4 as the foundation for its conductor model.
Base Model AgnosticEffect
significant step towards independence from specific foundational models for orchestration
From the article 2 mentionsThis efficiency allows for cost-effective retraining and experimentation with different base models.
Greater Customer ChoiceEffect
enables customers to select preferred base models, increasing sovereignty and flexibility
Modularity in ConductorsContext
conductor models can be swapped, allowing for diverse base model integration
From the article 9+ mentionsThis architecture is built upon research presented at ICLR 2026, involving Trinity and Conductor concepts.
Future OutlookEffect
continued exploration of various base models to enhance Fugu's adaptability
Contents(4)

Sakana AI has verified that its Fugu orchestration system can achieve comparable performance using Google's Gemma 4 as the foundation for its conductor model. This development marks a significant step towards a base model-agnostic approach for their multi-agent orchestration product.

Sakana Fugu's Two-Layer Architecture

Sakana Fugu functions as a unified platform for multi-agent orchestration. When a user submits a request, Fugu intelligently determines the processing path and dynamically calls upon a pool of high-performance models to generate a cohesive final answer. This architecture is built upon research presented at ICLR 2026, involving Trinity and Conductor concepts.

The system comprises two primary layers. The conductor model, a small language model trained by Sakana AI, dictates how other models should collaborate. The model pool consists of the actual processing units, which can include frontier models in the hundreds of billions of parameters. Crucially, the model pool is designed for flexibility, allowing components to be freely added or swapped.

Modularity in Conductor Models

The conductor model's role is not to possess all knowledge but to intelligently delegate tasks to specialized models. Since the heavy lifting of knowledge and inference is handled by the model pool, the conductor itself can remain small. This efficiency allows for cost-effective retraining and experimentation with different base models.

Sakana AI has consistently designed the model pool for interchangeability. While the standard configuration prioritizes top performance, the company recognizes varying customer needs. These can include cost optimization, geographical restrictions on model providers, or specific execution environment requirements. This flexibility was recently enhanced through a partnership with NVIDIA, enabling the use of Nemotron models.

Gemma 4 Verification

Until now, the diversity of base models for the conductor itself had not been thoroughly addressed. To rectify this, Sakana AI trained a conductor model using Gemma 4, an open model from a different lineage than their previous Qwen-based conductor. The objective was to confirm if the same training methodologies would prove effective across different model families of similar scale.

The evaluation utilized Sakana AI's proprietary benchmark suite, encompassing knowledge-based questions, code modification and generation, and graduate-level scientific inquiries. These questions were exclusively used for testing and were not part of the training or candidate selection process. Using Gemma 4 E2B as the conductor's base and applying Fugu's training regimen, the results showed performance on par with existing conductors. The cost-effectiveness was also found to be equivalent. A baseline of 'random distribution' represented the performance without specific training, where destinations were chosen arbitrarily from the model pool.

StartupHub.ai data indicates Sakana AI, which has raised $19M in 2024, holds a score of 70/100, placing it in a competitive space with peers like Adept AI (71/100) and Mistral AI (69/100).

Future Outlook

Previously, Sakana Fugu conductors were trained on Qwen. This Gemma 4 validation demonstrates that both the model pool and the conductor itself can be modularized, offering a wider array of choices. Sakana AI plans to train conductors based on proprietary models and prepare to offer conductors built on domestic Japanese models to meet customer sovereignty demands.

The company's ongoing development aims to balance access to global-leading AI capabilities with the assurance of required sovereignty. This strategic approach allows them to serve a diverse international clientele while respecting local regulations and preferences. Sakana AI is actively seeking talent to join their mission.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.