Harvey Labs: Building AI Research on a Budget

Harvey's Gabe Pereyra shares the playbook for building a competitive AI research lab on a budget, emphasizing benchmarks, synthetic data, and leveraging the frontier ecosystem.

Gabe Pereyra presenting at Sovereign AI event on "How Harvey Built a Research Lab on a Budget"
Sequoia Capital
Visual TL;DR
Frontier AI LabsDriver
billion-dollar funding, top talent, vast compute, and massive data advantages
From the article 6 mentionsIn a world where frontier AI labs command billions in funding, application-layer companies face a daunting challenge in building their own advanced AI capabilities.
Application-Layer ChallengeDriver
difficult for smaller companies to build advanced AI capabilities on a budget
From the article 4 mentionsIn a world where frontier AI labs command billions in funding, application-layer companies face a daunting challenge in building their own advanced AI capabilities.
Harvey's PlaybookCore
strategy for competitive AI research on a budget, shared by Gabe Pereyra
From the article 9+ mentionsHarvey, a company specializing in AI for legal and professional services, has developed a high-level playbook for establishing a research lab on a budget, as outlined by co-founder and president Gabe Pereyra.
Leverage EcosystemContext
strategically utilizing existing frontier ecosystem resources and open-source models
From the article 2 mentionsHowever, he emphasized that by strategically utilizing the existing "frontier ecosystem," companies can indeed compete and build "frontier intelligence."
Benchmarks & Post-TrainingCore
focus on specific benchmarks and post-training for production serving
From the article 6 mentionsBuild a benchmark: Pereyra stressed that without a solid benchmark, model training and subsequent production serving are impossible.
Synthetic DataCore
generating high-quality synthetic data with domain experts for training
From the article 4 mentionsThe breakthrough came through using domain experts to guide synthetic data generation.
Scaling with Neo LabsCore
partnering with Neo Labs and using open-source models for efficient scaling
From the article 2 mentionsHe recommended leveraging "neo labs", specialized teams or companies with expertise and infrastructure, to help with the post-training process.
Model Serving InfraCore
robust infrastructure for efficient model serving and continuous improvement
From the article 3 mentionsPereyra outlined Harvey's sophisticated model serving infrastructure, which handles multiple model families, fallbacks across providers to meet SLAs, and the integration of open-source models.
Competitive AIOutcome
building advanced AI capabilities despite budget constraints and resource gaps
Contents(7)

In a world where frontier AI labs command billions in funding, application-layer companies face a daunting challenge in building their own advanced AI capabilities. Harvey, a company specializing in AI for legal and professional services, has developed a high-level playbook for establishing a research lab on a budget, as outlined by co-founder and president Gabe Pereyra. Pereyra, who previously worked at DeepMind and Meta AI, shared Harvey's strategy for competing effectively in the rapidly evolving AI landscape.

Harvey Labs: Building AI Research on a Budget - Sequoia Capital
Harvey Labs: Building AI Research on a Budget, Sequoia Capital

The "Unfair Game" and Leveraging the Frontier Ecosystem

Pereyra opened by acknowledging the inherent disadvantage application-layer companies face against well-funded frontier labs, which possess greater financial resources, talent, compute infrastructure, and data. "There are rich teams, there are poor teams, and then there's us in the application layer," he stated. However, he emphasized that by strategically utilizing the existing "frontier ecosystem," companies can indeed compete and build "frontier intelligence."

Harvey's Playbook: Benchmarks, Post-Training, and Production Serving

Pereyra detailed Harvey's three-step playbook for building a research lab:

  • Build a benchmark: Pereyra stressed that without a solid benchmark, model training and subsequent production serving are impossible. Harvey has released several datasets, including Legal Agent Bench (LAB) for legal associate tasks, LAB Contracts for negotiation training, and LAB Diligence, an expansive RL environment designed for long-context, complex tasks.
  • Use it to post-train models: The core of this step involves using the custom benchmarks to refine open-source models.
  • Serve with inference providers: This final stage focuses on deploying these refined models into production environments.

Synthetic Data and Domain Experts

A significant challenge for Harvey, working with highly sensitive and privileged legal data from top law firms, is the inability to train on customer data. The breakthrough came through using domain experts to guide synthetic data generation. Pereyra likened it to how engineers now "vibe code" with coding models. Harvey's legal researchers, including Pereyra's own brother who is a lawyer at the firm, are trained to use AI tools to generate realistic datasets that mimic real-world scenarios. Companies like Mercor and Snorkel are then used to scale this process.

Scaling with Neo Labs and Open Source Models

Pereyra highlighted the increasing competitiveness of open-source models like Kimmi 3, GLM 5.2, and NeMo-Megatron. He recommended leveraging "neo labs", specialized teams or companies with expertise and infrastructure, to help with the post-training process. Harvey has partnered with various providers, including Fireworks, Base10, and Trajectory, to fine-tune these models for specific tasks. He noted that working with multiple neo labs allows Harvey to learn from different research approaches and model bets, and that the ease of post-training is rapidly increasing.

The Model Serving Infrastructure

Deploying models in production is a non-trivial task, especially for a company operating in 60 countries with diverse customer needs and model preferences. Pereyra outlined Harvey's sophisticated model serving infrastructure, which handles multiple model families, fallbacks across providers to meet SLAs, and the integration of open-source models. The decision to put a model into production, and to keep it there, is based on a combination of generic evaluations (like the LAB benchmark), human testing, critical user journeys, automated product tests, and heuristic signals like cost and latency.

The Post-Training Flywheel

Pereyra emphasized the importance of establishing a "post-training flywheel," where serving models in production and collecting feedback (without training on customer data) informs future dataset development and model improvements. He also touched upon the strategy of starting with "naive model swaps" and routing, where simpler tasks can be handled by open-source models, and then gradually incorporating more complex routing strategies.

The Future: Every Company as an AI Company

Pereyra concluded with a powerful analogy from Moneyball, suggesting that by winning on a budget and leveraging the frontier ecosystem, application-layer companies can fundamentally "change the game." He believes that in the future, every company will need to become an AI company and adopt similar playbooks. The key, he asserted, is to be strategic and utilize the available resources effectively.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.