Mercor's Brendan Foody on RL Environments for AI

Mercor's Brendan Foody details the shift to agentic data and RL environments, crucial for training frontier AI, and discusses the role of expert data and future trends.

Brendan Foody of Mercor presenting on RL Environments and Data.
Brendan Foody of Mercor discussing RL environments and AI data.· Sequoia Capital
Visual TL;DR
AI Data EvolutionContext
shifted from crowdsourced SFT/RLHF to agentic data for advanced models
From the article 9+ mentionsBrendan Foody of Mercor recently discussed the evolution of AI data and the critical role of Reinforcement Learning (RL) environments in training advanced AI models.
Agentic Data EraDriver
requires highly skilled experts collaborating to create frontier evaluations
From the article 4 mentionsThis era enabled significant progress in models like GPT-3, leading to advancements like ChatGPT, GPT-4, and other agentic data applications.
RL EnvironmentsCore
crucial for training frontier AI, enabling advanced agentic applications
From the article 9+ mentionsThis shift emphasizes the need for highly skilled experts who can collaborate to create frontier evaluations and RL environments.
Expert Data BottleneckDriver
lack of expert-generated data limits AI progress and model capabilities
Mercor's RoleCore
helps organizations build intelligence by focusing on post-training models
From the article 5 mentionsBrendan Foody of Mercor recently discussed the evolution of AI data and the critical role of Reinforcement Learning (RL) environments in training advanced AI models.
Future AI DataOutcome
focus on expert-driven, collaborative data for advanced AI systems
From the article 9+ mentionsMercor, a company that has seen substantial revenue growth, is at the forefront of helping organizations build their own intelligence by focusing on post-training models and agentic data.
Application LayerEffect
disseminating advanced AI capabilities to real-world applications
From the article 8 mentionsMercor's journey began with deep research projects, becoming a primary agentic data vendor for leading AI labs and application layer companies such as Harvey, Cera, and Cognition.
Contents(5)

Brendan Foody of Mercor recently discussed the evolution of AI data and the critical role of Reinforcement Learning (RL) environments in training advanced AI models. Mercor, a company that has seen substantial revenue growth, is at the forefront of helping organizations build their own intelligence by focusing on post-training models and agentic data.

Mercor's Brendan Foody on RL Environments for AI - Sequoia Capital
Mercor's Brendan Foody on RL Environments for AI, Sequoia Capital

The Evolution of Data for AI Training

Foody traced the history of data collection for AI, starting with crowdsourcing in 2020, which primarily involved supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) data. This era enabled significant progress in models like GPT-3, leading to advancements like ChatGPT, GPT-4, and other agentic data applications.

However, the market has transitioned towards an "agentic era" of data. This shift emphasizes the need for highly skilled experts who can collaborate to create frontier evaluations and RL environments. These experts, including software engineers, lawyers, and doctors, are essential for measuring and improving the capabilities of next-generation AI models.

Mercor's journey began with deep research projects, becoming a primary agentic data vendor for leading AI labs and application layer companies such as Harvey, Cera, and Cognition. Foody highlighted the recent evolution of RLVR (Reinforcement Learning from Human Preferences) to include RL environments, which feature rich applications and worlds designed to teach AI agents how to interact with everyday tools on a laptop.

Components of an RL Environment

Foody broke down the structure of an RL environment into three key parts:

  • Worlds: These encompass all the messages, slides, documents, and spreadsheets that mirror real-world projects and companies. They can be synthetically augmented or include acquired real-world data.
  • Apps: These are high-fidelity clones of popular applications like Salesforce, Microsoft 365, and ServiceNow, allowing agents to interact via command-line interfaces (CLI), MCP, or CUAs. Mocks enable state management for these applications.
  • Tasks: This component includes prompts and verifiers, which can be rubrics or unit tests used for evaluation or training.

The primary challenge for frontier AI labs is to cover the full distribution of these "worlds," "apps," and "tasks" across the economy. This requires an enormous scale-out, with humans playing a central role in building these environments, often with models in the loop.

Expert Data as the Bottleneck to AI Progress

The presentation included a graph illustrating the dramatic increase in expert hours per quarter, reaching 2.5 million hours in Q2 alone. This surge is driven by the need to scale environments across every economic category. Foody emphasized that while some domains, like mathematics, have clean simulation environments, most, such as creating slide decks, require human judgment for evaluation.

He used the analogy of a student grading their own homework, explaining that models struggle to reliably identify their own mistakes. This is where human experts create rubrics, similar to professors grading essays, to ensure accuracy and alignment with desired outcomes. Building these verifiers is complex, requiring an understanding of the problem space and potential errors.

Dissemination to the Application Layer

Foody noted that technology initially developed in frontier labs is now being disseminated to the application layer. Companies are realizing that their AI strategy hinges on three core pillars: compute, algorithms, and data sets, with data often being the most differentiating factor.

He showcased a sample legal environment, developed in collaboration with top law firms, that includes real-world scenarios and data rooms. These environments are then used to train AI models to perform specific tasks, such as evaluating legal liability. The process involves not only creating realistic data but also ensuring the accuracy of verifiers, a technically challenging task requiring extensive research and quality control.

The Future of AI Data and Mercor's Role

Foody highlighted recent advancements, noting that models like GLM 52 and Chimera K3 are now appearing on leaderboards, indicating progress in achieving frontier intelligence.

Mercor is actively working with customers like Harvey to build custom environments tailored to their specific domains, enabling them to develop proprietary AI capabilities. He cited Cursor as an early example of an application layer company successfully building an industry-leading model.

Regarding data curation, Foody outlined three primary methods:

  • By Task: Expert-authored evaluations priced per task, with companies potentially purchasing tens of thousands of tasks monthly.
  • Off-The-Shelf: Pre-built datasets that companies can purchase and run immediately, saving redundant development efforts across multiple labs.
  • Expert Staffing: Directing work and embedding vetted domain experts, allowing companies to scale their efforts as needed.

Foody concluded by expressing excitement about bringing advanced AI training technologies, previously exclusive to frontier labs, to a broader audience of application layer companies.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.