# Mercor's Brendan Foody on RL Environments for AI _Mercor's Brendan Foody details the shift to agentic data and RL environments, crucial for training frontier AI, and discusses the role of expert data and future trends._ **Published:** 2026-08-12 **Source:** https://www.startuphub.ai/ai-news/artificial-intelligence/2026/mercor-s-brendan-foody-on-rl-environments-for-ai --- Brendan Foody of Mercor recently discussed the evolution of AI data and the critical role of [Reinforcement Learning (RL) environments](/ai-news/ai/2026/coreweave-launches-ai-sandboxes) in training advanced AI models. Mercor, a company that has seen substantial revenue growth, is at the forefront of helping organizations build their own intelligence by focusing on post-training models and agentic data. AI Data EvolutionContext shifted from crowdsourced SFT/RLHF to agentic data for advanced modelsFrom the article 9+ mentionsBrendan Foody of Mercor recently discussed the evolution of AI data and the critical role of Reinforcement Learning (RL) environments in training advanced AI models.leads toAgentic Data EraDriverrequires highly skilled experts collaborating to create frontier evaluationsFrom the article 4 mentionsThis era enabled significant progress in models like GPT-3, leading to advancements like ChatGPT, GPT-4, and other agentic data applications.needsRL EnvironmentsCorecrucial for training frontier AI, enabling advanced agentic applicationsFrom the article 9+ mentionsThis shift emphasizes the need for highly skilled experts who can collaborate to create frontier evaluations and RL environments.limited byExpert Data BottleneckDriverlack of expert-generated data limits AI progress and model capabilitiesaddressed byMercor's RoleCorehelps organizations build intelligence by focusing on post-training modelsFrom the article 5 mentionsBrendan Foody of Mercor recently discussed the evolution of AI data and the critical role of Reinforcement Learning (RL) environments in training advanced AI models.shapesFuture AI DataOutcomefocus on expert-driven, collaborative data for advanced AI systemsFrom the article 9+ mentionsMercor, a company that has seen substantial revenue growth, is at the forefront of helping organizations build their own intelligence by focusing on post-training models and agentic data.informsApplication LayerEffectdisseminating advanced AI capabilities to real-world applicationsFrom the article 8 mentionsMercor's journey began with deep research projects, becoming a primary agentic data vendor for leading AI labs and application layer companies such as Harvey, Cera, and Cognition. ## The Evolution of Data for AI Training Foody traced the history of data collection for AI, starting with crowdsourcing in 2020, which primarily involved supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) data. This era enabled significant progress in models like GPT-3, leading to advancements like ChatGPT, GPT-4, and other agentic data applications. However, the market has transitioned towards an "agentic era" of data. This shift emphasizes the need for highly skilled experts who can collaborate to create frontier evaluations and RL environments. These experts, including software engineers, lawyers, and doctors, are essential for measuring and improving the capabilities of next-generation AI models. Mercor's journey began with deep research projects, becoming a primary agentic data vendor for leading AI labs and application layer companies such as Harvey, Cera, and Cognition. Foody highlighted the recent evolution of RLVR (Reinforcement Learning from Human Preferences) to include RL environments, which feature rich applications and worlds designed to teach AI agents how to interact with everyday tools on a laptop. ## Components of an RL Environment Foody broke down the structure of an RL environment into three key parts: - **Worlds:** These encompass all the messages, slides, documents, and spreadsheets that mirror real-world projects and companies. They can be synthetically augmented or include acquired real-world data. - **Apps:** These are high-fidelity clones of popular applications like Salesforce, Microsoft 365, and ServiceNow, allowing agents to interact via command-line interfaces (CLI), MCP, or CUAs. Mocks enable state management for these applications. - **Tasks:** This component includes prompts and verifiers, which can be rubrics or unit tests used for evaluation or training. The primary challenge for frontier AI labs is to cover the full distribution of these "worlds," "apps," and "tasks" across the economy. This requires an enormous scale-out, with humans playing a central role in building these environments, often with models in the loop. ## Expert Data as the Bottleneck to AI Progress The presentation included a graph illustrating the dramatic increase in expert hours per quarter, reaching 2.5 million hours in Q2 alone. This surge is driven by the need to scale environments across every economic category. Foody emphasized that while some domains, like mathematics, have clean simulation environments, most, such as creating slide decks, require human judgment for evaluation. He used the analogy of a student grading their own homework, explaining that models struggle to reliably identify their own mistakes. This is where human experts create rubrics, similar to professors grading essays, to ensure accuracy and alignment with desired outcomes. Building these verifiers is complex, requiring an understanding of the problem space and potential errors. ## Dissemination to the Application Layer Foody noted that technology initially developed in frontier labs is now being disseminated to the application layer. Companies are realizing that their AI strategy hinges on three core pillars: compute, algorithms, and data sets, with data often being the most differentiating factor. He showcased a sample legal environment, developed in collaboration with top law firms, that includes real-world scenarios and data rooms. These environments are then used to train AI models to perform specific tasks, such as evaluating legal liability. The process involves not only creating realistic data but also ensuring the accuracy of verifiers, a technically challenging task requiring extensive research and quality control. ## The Future of AI Data and Mercor's Role Foody highlighted recent advancements, noting that models like GLM 52 and Chimera K3 are now appearing on leaderboards, indicating progress in achieving frontier intelligence. Mercor is actively working with customers like Harvey to build custom environments tailored to their specific domains, enabling them to develop proprietary AI capabilities. He cited Cursor as an early example of an application layer company successfully building an industry-leading model. Regarding data curation, Foody outlined three primary methods: - **By Task:** Expert-authored evaluations priced per task, with companies potentially purchasing tens of thousands of tasks monthly. - **Off-The-Shelf:** Pre-built datasets that companies can purchase and run immediately, saving redundant development efforts across multiple labs. - **Expert Staffing:** Directing work and embedding vetted domain experts, allowing companies to scale their efforts as needed. Foody concluded by expressing excitement about bringing advanced AI training technologies, previously exclusive to frontier labs, to a broader audience of application layer companies. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.