# LangChain CEO on Building Better AI Agents _LangChain CEO Harrison Chase discusses the essential components of AI agents, the importance of custom harnesses, and the power of evals and observability for continuous improvement._ **Published:** 2026-08-13 **Source:** https://www.startuphub.ai/ai-news/artificial-intelligence/2026/langchain-ceo-on-building-better-ai-agents --- Harrison Chase, co-founder and CEO of LangChain, recently shared his insights on building and refining AI agents, focusing on the critical roles of "harnesses" and "evals" in developing truly intelligent systems. Speaking at a Sovereign AI event, Chase emphasized that to truly "own your intelligence," developers need to have control over the three core components of an agent: the model, the context, and the harness. AI Agent ComponentsContext From the article 9+ mentionsSpeaking at a Sovereign AI event, Chase emphasized that to truly "own your intelligence," developers need to have control over the three core components of an agent: the model, the context, and the harness.Own IntelligenceEffectFrom the article 2 mentionsSpeaking at a Sovereign AI event, Chase emphasized that to truly "own your intelligence," developers need to have control over the three core components of an agent: the model, the context, and the harness.The HarnessCoreorchestrates agent behavior, central to building intelligent systemsFrom the article 8 mentionsFinally, the harness, which Chase highlighted as the central focus, is responsible for orchestrating the model and context.requiresCustom HarnessesEffectcrucial for avoiding lock-in and leveraging best available technologyFrom the article 6 mentionsChase addressed the common question of when to build a custom harness versus using an off-the-shelf solution.drivesEvals & ObservabilityCorepower continuous improvement and refine agent behavior over timeFrom the article 7 mentionsThe second major pillar of Chase's talk was the importance of evals and observability.createsData FlywheelContextevals and observability create a feedback loop for automated improvementFrom the article 2 mentionsChase concluded by outlining a "recipe for continuously improving agents": build a v1 agent, collect traces, curate trace data, and run experiments on that data.leads toAutomate ImprovementOutcomeleverage data flywheel to automatically enhance agent performance and intelligenceFrom the article 2 mentionsHe also highlighted the potential to automate this process, introducing LangSmith Engine, an agent built on top of traces that can perform data curation and suggest fixes. ## Understanding Agent Components Chase broke down an agent into three main parts. First, the **model** itself, which can be an open-weight model or a proprietary one, with the ability to switch models being crucial for avoiding lock-in and leveraging the best available technology. Second, the **context**, which includes memory, semantic knowledge retrieved through methods like RAG, and previous conversational history, all of which help personalize and guide the agent's behavior. Finally, the **harness**, which Chase highlighted as the central focus, is responsible for orchestrating the model and context. The harness's primary job is to deliver the right context to the model at the precise moment it's needed, enabling the agent to perform complex tasks. ## The Role of the Harness The harness acts as the conductor, managing the flow of information and ensuring the agent can interact with external systems, process their outputs, and feed them back into the loop. At its simplest, an agent is an LLM running in a loop, calling tools. However, Chase explained that harnesses allow for significant customization. He illustrated this with LangChain's own minimal harness and compared it to "Deep Agents," a more generalized version. Customizations can be made through "middleware constructs" or "hooks and plugins" that modify the core loop. These allow for actions like running code snippets before agent invocation, wrapping model or tool calls, and managing context through summarization or offloading. These elements enable agents to access sandboxes, file systems, memory, and more, tailoring them to specific domains. ## Customization vs. Off-the-Shelf Harnesses Chase addressed the common question of when to build a custom harness versus using an off-the-shelf solution. He noted that general-purpose harnesses are often sufficient for basic tasks, especially when starting. However, as use cases become more specialized or "out-of-distribution," customizing the harness becomes more important. He used the example of "legal AI," where specific tasks like editing files are well-represented in model training, but the broader task requires a custom harness. LangChain's "Deep Agents" utilizes "model profiles" to switch between different implementations of tools like "edit file" based on the model being used, demonstrating a pragmatic approach to customization. ## Evals and Observability: The Data Flywheel The second major pillar of Chase's talk was the importance of **evals and observability**. He quoted Satya Nadella, emphasizing the need to "create your private evals, because evals define what 'good' looks like inside the organization." This includes retaining ownership of organizational memory, traces, feedback, decisions, and institutional context. Evals are crucial for defining benchmarks, catching regressions, and "hill climbing" on performance by adjusting models or harnesses. Chase highlighted **Harbor**, an open-source eval runner, as an emerging industry standard. Harbor allows users to run agents against datasets, with each task defined by an environment (e.g., a Dockerfile), an instruction (in Markdown), and an evaluation script. The results can then be visualized and compared, providing valuable insights into agent performance. He also touched upon observability, stressing its underrated importance for debugging agents. When agents fail, it's often due to issues with the model or, more commonly, the context provided. Detailed observability into the context window, step execution, and tool usage is vital for effective debugging. ## Automating Improvement Chase concluded by outlining a "recipe for continuously improving agents": build a v1 agent, collect traces, curate trace data, and run experiments on that data. He also highlighted the potential to automate this process, introducing LangSmith Engine, an agent built on top of traces that can perform data curation and suggest fixes. This automated approach aims to streamline the development cycle and accelerate agent improvement, ultimately leading to more capable and reliable AI systems. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.