Trajectory's Ronak Malde on Scaling Continual Learning

Trajectory founder Ronak Malde discusses the limitations of current AI scaling and introduces On-Policy Self-Distillation (OPSD) as a solution for continual learning.

8 min read
Ronak Malde speaking on stage at AI Engineer World's Fair
AI Engineer

Visual TL;DR. Benchmark Bottleneck identifies Trajectory's Ronak Malde. Limitations of Algorithms also notes Trajectory's Ronak Malde. Trajectory's Ronak Malde proposes On-Policy Self-Distillation. On-Policy Self-Distillation faces Scaling Challenges. On-Policy Self-Distillation enables Continual Learning. Continual Learning unlocks Trillion-Token Opportunity. Continual Learning leads to New AI Development.

  1. Benchmark Bottleneck: current AI scaling relies on costly, time-consuming benchmarks not tied to real-world use
  2. Trajectory's Ronak Malde: founder of Trajectory, presenting a new approach to AI continual learning
  3. Limitations of Algorithms: existing AI algorithms struggle with continuous learning from real-world interactions
  4. On-Policy Self-Distillation: OPSD introduced as a solution for efficient and continuous AI learning
  5. Scaling Challenges: addressing the difficulties in implementing OPSD for large-scale AI systems
  6. Continual Learning: enabling AI to learn continuously from real-world interactions, like humans
  7. Trillion-Token Opportunity: future AI will learn from vast amounts of real-world data, not just benchmarks
  8. New AI Development: shifting AI development towards adaptive, real-world learning paradigms
Visual TL;DR
Visual TL;DR, startuphub.ai Benchmark Bottleneck identifies Trajectory's Ronak Malde. Trajectory's Ronak Malde proposes On-Policy Self-Distillation. On-Policy Self-Distillation enables Continual Learning. Continual Learning leads to New AI Development identifies proposes enables leads to Benchmark Bottleneck Trajectory's Ronak Malde On-Policy Self-Distillation Continual Learning New AI Development From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Benchmark Bottleneck identifies Trajectory's Ronak Malde. Trajectory's Ronak Malde proposes On-Policy Self-Distillation. On-Policy Self-Distillation enables Continual Learning. Continual Learning leads to New AI Development identifies proposes enables leads to BenchmarkBottleneck Trajectory'sRonak Malde On-PolicySelf-Distillation ContinualLearning New AIDevelopment From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Benchmark Bottleneck identifies Trajectory's Ronak Malde. Trajectory's Ronak Malde proposes On-Policy Self-Distillation. On-Policy Self-Distillation enables Continual Learning. Continual Learning leads to New AI Development identifies proposes enables leads to Benchmark Bottleneck current AI scaling relies on costly,time-consuming benchmarks not tied toreal-world use Trajectory's Ronak Malde founder of Trajectory, presenting a newapproach to AI continual learning On-Policy Self-Distillation OPSD introduced as a solution forefficient and continuous AI learning Continual Learning enabling AI to learn continuously fromreal-world interactions, like humans New AI Development shifting AI development towards adaptive,real-world learning paradigms From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Benchmark Bottleneck identifies Trajectory's Ronak Malde. Trajectory's Ronak Malde proposes On-Policy Self-Distillation. On-Policy Self-Distillation enables Continual Learning. Continual Learning leads to New AI Development identifies proposes enables leads to BenchmarkBottleneck current AI scalingrelies on costly,time-consuming… Trajectory'sRonak Malde founder ofTrajectory,presenting a new… On-PolicySelf-Distillation OPSD introduced asa solution forefficient and… ContinualLearning enabling AI tolearn continuouslyfrom real-world… New AIDevelopment shifting AIdevelopment towardsadaptive,… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Benchmark Bottleneck identifies Trajectory's Ronak Malde. Limitations of Algorithms also notes Trajectory's Ronak Malde. Trajectory's Ronak Malde proposes On-Policy Self-Distillation. On-Policy Self-Distillation faces Scaling Challenges. On-Policy Self-Distillation enables Continual Learning. Continual Learning unlocks Trillion-Token Opportunity. Continual Learning leads to New AI Development identifies also notes proposes faces enables unlocks leads to Benchmark Bottleneck current AI scaling relies on costly,time-consuming benchmarks not tied toreal-world use Trajectory's Ronak Malde founder of Trajectory, presenting a newapproach to AI continual learning Limitations of Algorithms existing AI algorithms struggle withcontinuous learning from real-worldinteractions On-Policy Self-Distillation OPSD introduced as a solution forefficient and continuous AI learning Scaling Challenges addressing the difficulties inimplementing OPSD for large-scale AIsystems Continual Learning enabling AI to learn continuously fromreal-world interactions, like humans Trillion-Token Opportunity future AI will learn from vast amounts ofreal-world data, not just benchmarks New AI Development shifting AI development towards adaptive,real-world learning paradigms From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Benchmark Bottleneck identifies Trajectory's Ronak Malde. Limitations of Algorithms also notes Trajectory's Ronak Malde. Trajectory's Ronak Malde proposes On-Policy Self-Distillation. On-Policy Self-Distillation faces Scaling Challenges. On-Policy Self-Distillation enables Continual Learning. Continual Learning unlocks Trillion-Token Opportunity. Continual Learning leads to New AI Development identifies also notes proposes faces enables unlocks leads to BenchmarkBottleneck current AI scalingrelies on costly,time-consuming… Trajectory'sRonak Malde founder ofTrajectory,presenting a new… Limitations ofAlgorithms existing AIalgorithms strugglewith continuous… On-PolicySelf-Distillation OPSD introduced asa solution forefficient and… ScalingChallenges addressing thedifficulties inimplementing OPSD… ContinualLearning enabling AI tolearn continuouslyfrom real-world… Trillion-TokenOpportunity future AI willlearn from vastamounts of… New AIDevelopment shifting AIdevelopment towardsadaptive,… From startuphub.ai · The publishers behind this format

Ronak Malde, founder of Trajectory, a platform focused on continual learning for AI, presented a compelling case for a new approach to AI development. Speaking at the AI Engineer World's Fair, Malde outlined the limitations of current AI scaling methods, which rely heavily on benchmarks that quickly saturate and become costly to maintain. He argued that the future of AI lies in its ability to learn continuously from real-world interactions, much like humans do.

Trajectory's Ronak Malde on Scaling Continual Learning - AI Engineer
Trajectory's Ronak Malde on Scaling Continual Learning — from AI Engineer

The Benchmark Bottleneck

Malde began by illustrating the trend of rapidly scaling AI benchmarks over the past few years. While this has led to significant progress, he noted that these benchmarks are becoming increasingly time-consuming and expensive to train on. "We're seeing domains where it takes 4 hours, 6 hours, 24 hours or even several days in order to scale up benchmarks," Malde stated. More concerningly, he pointed out that these benchmarks are often "not tied to real world use cases where people are using AI."

The Trillion-Token Opportunity

In contrast, Malde highlighted the vast opportunity presented by the trillions of tokens generated daily through AI inference. This real-world data, capturing how models perform both well and poorly, should be the signal used for continuous learning. "This is actually how humans learn, right? We're continuously updating in the real world and getting smarter every single day," he explained. The field is increasingly recognizing this, with prominent AI experts and industry leaders like Ilya Sutskever, Andrej Karpathy, Satya Nadella, and Demis Hassabis all emphasizing the importance of continual learning.

Limitations of Current Algorithms

Malde then delved into the core problems hindering the widespread adoption of effective continual learning. He identified four key issues with current algorithms:

  • Task distribution mismatch: Benchmarks often don't reflect real-world scenarios.
  • Off-policy sampling: Methods learn from behavior that the model didn't actually generate.
  • Rollout parallelism: Training requires massive infrastructure to replicate real-world environments, introducing bias.
  • Sequence-level reward: A single score for an entire trajectory misses nuanced, per-token feedback.

He charted the evolution of these methods, from SFT (Instruction Fine-tuning) to DPO (Direct Preference Optimization) and GRPO (Generalized Reinforcement Learning Policy Optimization), noting that while each iteration has improved, none have fully solved these challenges.

Introducing On-Policy Self-Distillation (OPSD)

Malde introduced On-Policy Self-Distillation (OPSD) as a novel approach designed to overcome these limitations. The core idea is to use a model's own rollouts as the training data, guided by a 'teacher' model. In the case of self-distillation, when a superior teacher model isn't available, the model can leverage 'privileged information' or hints within its prompts to become its own teacher.

"You take what's called this hint, put it into the beginning of the prompt, and now you match the log props of the student without that hint to the teacher with that hint," Malde explained. This method, he asserted, "is an extremely powerful algorithm" that addresses the four key problems: it uses online task distributions, ensures on-policy sampling, requires only single parallel rollouts, and provides dense, per-token reward signals.

Scaling Challenges and Solutions

While OPSD shows great promise, Malde acknowledged the challenges encountered when scaling it up to larger models and longer tasks. He highlighted the "but wait" problem, where a model might get stuck in suboptimal loops, and the issue of "hint leakage," where privileged information might prematurely reveal the answer. To address these, Trajectory is developing techniques like step-level divergence weighting and residual guidance.

These advancements have allowed them to achieve significant improvements, as demonstrated by their work on a 12B parameter model for Mercor Apex Agents, which surpassed traditional RL methods. Malde concluded by emphasizing Trajectory's vision: "Agents that get better the more they're used." The company is building a platform to facilitate this continuous learning loop, turning every interaction into an opportunity for improvement.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.