Mercor's Brendan Foody on RL Environments for AI

Mercor's Brendan Foody details the shift to agentic data and RL environments, crucial for training frontier AI, and discusses the role of expert data and future trends.

8 min read
Brendan Foody of Mercor presenting on RL Environments and Data.
Brendan Foody of Mercor discussing RL environments and AI data.· Sequoia Capital

Visual TL;DR. AI Data Evolution leads to Agentic Data Era. Agentic Data Era needs RL Environments. RL Environments limited by Expert Data Bottleneck. Expert Data Bottleneck addressed by Mercor's Role. Mercor's Role shapes Future AI Data. RL Environments enables Application Layer. Future AI Data informs Application Layer.

  1. AI Data Evolution: shifted from crowdsourced SFT/RLHF to agentic data for advanced models
  2. Agentic Data Era: requires highly skilled experts collaborating to create frontier evaluations
  3. RL Environments: crucial for training frontier AI, enabling advanced agentic applications
  4. Expert Data Bottleneck: lack of expert-generated data limits AI progress and model capabilities
  5. Mercor's Role: helps organizations build intelligence by focusing on post-training models
  6. Future AI Data: focus on expert-driven, collaborative data for advanced AI systems
  7. Application Layer: disseminating advanced AI capabilities to real-world applications
Visual TL;DR
Visual TL;DR, startuphub.ai RL Environments limited by Expert Data Bottleneck. Expert Data Bottleneck addressed by Mercor's Role limited by addressed by AI Data Evolution RL Environments Expert Data Bottleneck Mercor's Role From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai RL Environments limited by Expert Data Bottleneck. Expert Data Bottleneck addressed by Mercor's Role limited by addressed by AI Data Evolution RL Environments Expert DataBottleneck Mercor's Role From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai RL Environments limited by Expert Data Bottleneck. Expert Data Bottleneck addressed by Mercor's Role limited by addressed by AI Data Evolution shifted from crowdsourced SFT/RLHF toagentic data for advanced models RL Environments crucial for training frontier AI, enablingadvanced agentic applications Expert Data Bottleneck lack of expert-generated data limits AIprogress and model capabilities Mercor's Role helps organizations build intelligence byfocusing on post-training models From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai RL Environments limited by Expert Data Bottleneck. Expert Data Bottleneck addressed by Mercor's Role limited by addressed by AI Data Evolution shifted fromcrowdsourcedSFT/RLHF to agentic… RL Environments crucial fortraining frontierAI, enabling… Expert DataBottleneck lack ofexpert-generateddata limits AI… Mercor's Role helps organizationsbuild intelligenceby focusing on… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Data Evolution leads to Agentic Data Era. Agentic Data Era needs RL Environments. RL Environments limited by Expert Data Bottleneck. Expert Data Bottleneck addressed by Mercor's Role. Mercor's Role shapes Future AI Data. RL Environments enables Application Layer. Future AI Data informs Application Layer leads to needs limited by addressed by shapes enables informs AI Data Evolution shifted from crowdsourced SFT/RLHF toagentic data for advanced models Agentic Data Era requires highly skilled expertscollaborating to create frontierevaluations RL Environments crucial for training frontier AI, enablingadvanced agentic applications Expert Data Bottleneck lack of expert-generated data limits AIprogress and model capabilities Mercor's Role helps organizations build intelligence byfocusing on post-training models Future AI Data focus on expert-driven, collaborative datafor advanced AI systems Application Layer disseminating advanced AI capabilities toreal-world applications From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Data Evolution leads to Agentic Data Era. Agentic Data Era needs RL Environments. RL Environments limited by Expert Data Bottleneck. Expert Data Bottleneck addressed by Mercor's Role. Mercor's Role shapes Future AI Data. RL Environments enables Application Layer. Future AI Data informs Application Layer leads to needs limited by addressed by shapes enables informs AI Data Evolution shifted fromcrowdsourcedSFT/RLHF to agentic… Agentic Data Era requires highlyskilled expertscollaborating to… RL Environments crucial fortraining frontierAI, enabling… Expert DataBottleneck lack ofexpert-generateddata limits AI… Mercor's Role helps organizationsbuild intelligenceby focusing on… Future AI Data focus onexpert-driven,collaborative data… Application Layer disseminatingadvanced AIcapabilities to… From startuphub.ai · The publishers behind this format

Brendan Foody of Mercor recently discussed the evolution of AI data and the critical role of Reinforcement Learning (RL) environments in training advanced AI models. Mercor, a company that has seen substantial revenue growth, is at the forefront of helping organizations build their own intelligence by focusing on post-training models and agentic data.

Mercor's Brendan Foody on RL Environments for AI - Sequoia Capital
Mercor's Brendan Foody on RL Environments for AI — from Sequoia Capital

The Evolution of Data for AI Training

Foody traced the history of data collection for AI, starting with crowdsourcing in 2020, which primarily involved supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) data. This era enabled significant progress in models like GPT-3, leading to advancements like ChatGPT, GPT-4, and other agentic data applications.

However, the market has transitioned towards an "agentic era" of data. This shift emphasizes the need for highly skilled experts who can collaborate to create frontier evaluations and RL environments. These experts, including software engineers, lawyers, and doctors, are essential for measuring and improving the capabilities of next-generation AI models.

Mercor's journey began with deep research projects, becoming a primary agentic data vendor for leading AI labs and application layer companies such as Harvey, Cera, and Cognition. Foody highlighted the recent evolution of RLVR (Reinforcement Learning from Human Preferences) to include RL environments, which feature rich applications and worlds designed to teach AI agents how to interact with everyday tools on a laptop.

Components of an RL Environment

Foody broke down the structure of an RL environment into three key parts:

  • Worlds: These encompass all the messages, slides, documents, and spreadsheets that mirror real-world projects and companies. They can be synthetically augmented or include acquired real-world data.
  • Apps: These are high-fidelity clones of popular applications like Salesforce, Microsoft 365, and ServiceNow, allowing agents to interact via command-line interfaces (CLI), MCP, or CUAs. Mocks enable state management for these applications.
  • Tasks: This component includes prompts and verifiers, which can be rubrics or unit tests used for evaluation or training.

The primary challenge for frontier AI labs is to cover the full distribution of these "worlds," "apps," and "tasks" across the economy. This requires an enormous scale-out, with humans playing a central role in building these environments, often with models in the loop.

Expert Data as the Bottleneck to AI Progress

The presentation included a graph illustrating the dramatic increase in expert hours per quarter, reaching 2.5 million hours in Q2 alone. This surge is driven by the need to scale environments across every economic category. Foody emphasized that while some domains, like mathematics, have clean simulation environments, most, such as creating slide decks, require human judgment for evaluation.

He used the analogy of a student grading their own homework, explaining that models struggle to reliably identify their own mistakes. This is where human experts create rubrics, similar to professors grading essays, to ensure accuracy and alignment with desired outcomes. Building these verifiers is complex, requiring an understanding of the problem space and potential errors.

Dissemination to the Application Layer

Foody noted that technology initially developed in frontier labs is now being disseminated to the application layer. Companies are realizing that their AI strategy hinges on three core pillars: compute, algorithms, and data sets, with data often being the most differentiating factor.

He showcased a sample legal environment, developed in collaboration with top law firms, that includes real-world scenarios and data rooms. These environments are then used to train AI models to perform specific tasks, such as evaluating legal liability. The process involves not only creating realistic data but also ensuring the accuracy of verifiers, a technically challenging task requiring extensive research and quality control.

The Future of AI Data and Mercor's Role

Foody highlighted recent advancements, noting that models like GLM 52 and Chimera K3 are now appearing on leaderboards, indicating progress in achieving frontier intelligence.

Mercor is actively working with customers like Harvey to build custom environments tailored to their specific domains, enabling them to develop proprietary AI capabilities. He cited Cursor as an early example of an application layer company successfully building an industry-leading model.

Regarding data curation, Foody outlined three primary methods:

  • By Task: Expert-authored evaluations priced per task, with companies potentially purchasing tens of thousands of tasks monthly.
  • Off-The-Shelf: Pre-built datasets that companies can purchase and run immediately, saving redundant development efforts across multiple labs.
  • Expert Staffing: Directing work and embedding vetted domain experts, allowing companies to scale their efforts as needed.

Foody concluded by expressing excitement about bringing advanced AI training technologies, previously exclusive to frontier labs, to a broader audience of application layer companies.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.