The AI race for raw reasoning power has hit a practical limit. Users demand frontier-level intelligence, the kind that tackles complex math problems, but not at the cost of minutes-long waits or exorbitant API fees. Enter Step 3.5 Flash, a new model from the StepFun Team designed to bridge this gap.
Companies working on this
Profiles of the companies named in this story, with founding year, headquarters, and a short description from our database.
Gen Z-focused fintech app offering financial services for credit building, saving, and investing.
- Founded
- 2018
- Location
- San Francisco, United States
- Valuation
- $1.0B
Intelligence Density Over Brute Force
Step 3.5 Flash operates on a principle of "High-Density Intelligence." It pairs a massive 196 billion parameter foundation with a highly efficient 11 billion active parameter execution engine. This architecture allows it to compete with models like GPT-5.2 xHigh and Gemini 3.0 Pro while maintaining agility.
The Architecture: Smarter, Not Just Bigger
At its core, Step 3.5 Flash utilizes a Sparse Mixture-of-Experts (MoE) backbone. While the total parameter count is 196B, only 11B are engaged per token. This drastically reduces computational overhead.
Key innovations include:
- Hybrid Attention: A 3:1 ratio of Sliding Window Attention (SWA) and Full Attention balances local detail capture with long-range dependency understanding.
- Multi-Token Prediction (MTP-3): The model generates three tokens simultaneously, significantly accelerating output for tasks requiring long sequences.
- Balanced Routing: EP-Group Balanced Routing ensures even hardware utilization, preventing bottlenecks and expert collapse, a common issue in Sparse MoE AI development.
Training for Stability and Scale
Forged over 17.2 trillion tokens, Step 3.5 Flash prioritizes training stability. The StepFun Team developed an asynchronous metrics server to monitor training at the micro-batch level, proactively identifying and mitigating precision issues and expert collapse. This meticulous approach ensures reliability in large-scale training.
