Dmitry Petrov: Bridging Agents and Physical Data

Dmitry Petrov of DataChain explains how specialized 'data harnesses' are crucial for enabling AI agents to effectively process unstructured physical data like videos and sensor logs.

9 min read
Dmitry Petrov, Co-Founder of DataChain, presenting on AI agents and physical data.
AI Engineer

Visual TL;DR. AI Agent Data Disconnect leads to Low AI Accuracy. AI Agent Data Disconnect addressed by DataChain's Petrov. Low AI Accuracy requires Agent Data Harnesses. DataChain's Petrov proposes Agent Data Harnesses. Agent Data Harnesses involves Overcoming Data Spaghetti. Overcoming Data Spaghetti uses Pydantic & Unified Stack. Agent Data Harnesses example DataChain in Action. DataChain in Action achieves Effective Agent Processing. Pydantic & Unified Stack enables Effective Agent Processing. Agent Data Harnesses enables Effective Agent Processing.

  1. AI Agent Data Disconnect: agents struggle with unstructured physical data like video, sensor logs, and robot data
  2. Low AI Accuracy: research shows 21% accuracy without specific data harnesses and context for agents
  3. DataChain's Petrov: Dmitry Petrov addresses the critical gap in processing messy real-world physical data
  4. Agent Data Harnesses: specialized 'data harnesses' are crucial for agents to process unstructured physical data
  5. Overcoming Data Spaghetti: moving beyond simple JSON to structured databases for complex physical data
  6. Pydantic & Unified Stack: leveraging Pydantic for data validation and a unified stack for agent data processing
  7. DataChain in Action: analyzing dashcam footage demonstrates the practical application of data harnesses
  8. Effective Agent Processing: enabling AI agents to effectively process and understand unstructured physical data
Visual TL;DR
Visual TL;DR, startuphub.ai AI Agent Data Disconnect addressed by DataChain's Petrov. DataChain's Petrov proposes Agent Data Harnesses. Agent Data Harnesses enables Effective Agent Processing addressed by proposes enables AI Agent Data Disconnect DataChain's Petrov Agent Data Harnesses Effective Agent Processing From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Agent Data Disconnect addressed by DataChain's Petrov. DataChain's Petrov proposes Agent Data Harnesses. Agent Data Harnesses enables Effective Agent Processing addressed by proposes enables AI Agent DataDisconnect DataChain'sPetrov Agent DataHarnesses Effective AgentProcessing From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Agent Data Disconnect addressed by DataChain's Petrov. DataChain's Petrov proposes Agent Data Harnesses. Agent Data Harnesses enables Effective Agent Processing addressed by proposes enables AI Agent Data Disconnect agents struggle with unstructured physicaldata like video, sensor logs, and robotdata DataChain's Petrov Dmitry Petrov addresses the critical gapin processing messy real-world physicaldata Agent Data Harnesses specialized 'data harnesses' are crucialfor agents to process unstructuredphysical data Effective Agent Processing enabling AI agents to effectively processand understand unstructured physical data From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Agent Data Disconnect addressed by DataChain's Petrov. DataChain's Petrov proposes Agent Data Harnesses. Agent Data Harnesses enables Effective Agent Processing addressed by proposes enables AI Agent DataDisconnect agents strugglewith unstructuredphysical data like… DataChain'sPetrov Dmitry Petrovaddresses thecritical gap in… Agent DataHarnesses specialized 'dataharnesses' arecrucial for agents… Effective AgentProcessing enabling AI agentsto effectivelyprocess and… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Agent Data Disconnect leads to Low AI Accuracy. AI Agent Data Disconnect addressed by DataChain's Petrov. Low AI Accuracy requires Agent Data Harnesses. DataChain's Petrov proposes Agent Data Harnesses. Agent Data Harnesses involves Overcoming Data Spaghetti. Overcoming Data Spaghetti uses Pydantic & Unified Stack. Agent Data Harnesses example DataChain in Action. DataChain in Action achieves Effective Agent Processing. Pydantic & Unified Stack enables Effective Agent Processing. Agent Data Harnesses enables Effective Agent Processing leads to addressed by requires proposes involves uses example achieves enables enables AI Agent Data Disconnect agents struggle with unstructured physicaldata like video, sensor logs, and robotdata Low AI Accuracy research shows 21% accuracy withoutspecific data harnesses and context foragents DataChain's Petrov Dmitry Petrov addresses the critical gapin processing messy real-world physicaldata Agent Data Harnesses specialized 'data harnesses' are crucialfor agents to process unstructuredphysical data Overcoming Data Spaghetti moving beyond simple JSON to structureddatabases for complex physical data Pydantic & Unified Stack leveraging Pydantic for data validationand a unified stack for agent dataprocessing DataChain in Action analyzing dashcam footage demonstrates thepractical application of data harnesses Effective Agent Processing enabling AI agents to effectively processand understand unstructured physical data From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai AI Agent Data Disconnect leads to Low AI Accuracy. AI Agent Data Disconnect addressed by DataChain's Petrov. Low AI Accuracy requires Agent Data Harnesses. DataChain's Petrov proposes Agent Data Harnesses. Agent Data Harnesses involves Overcoming Data Spaghetti. Overcoming Data Spaghetti uses Pydantic & Unified Stack. Agent Data Harnesses example DataChain in Action. DataChain in Action achieves Effective Agent Processing. Pydantic & Unified Stack enables Effective Agent Processing. Agent Data Harnesses enables Effective Agent Processing leads to addressed by requires proposes involves uses example achieves enables enables AI Agent DataDisconnect agents strugglewith unstructuredphysical data like… Low AI Accuracy research shows 21%accuracy withoutspecific data… DataChain'sPetrov Dmitry Petrovaddresses thecritical gap in… Agent DataHarnesses specialized 'dataharnesses' arecrucial for agents… Overcoming DataSpaghetti moving beyondsimple JSON tostructured… Pydantic &Unified Stack leveraging Pydanticfor data validationand a unified stack… DataChain inAction analyzing dashcamfootagedemonstrates the… Effective AgentProcessing enabling AI agentsto effectivelyprocess and… From startuphub.ai · The publishers behind this format

In the rapidly evolving world of AI, the ability for agents to effectively process and understand unstructured physical data remains a significant challenge. Dmitry Petrov, co-founder of DataChain, recently addressed this critical gap in his presentation, "When Agents Meet Physical Data: The Other Physics of Agent Harnesses." Petrov highlighted the stark contrast between the current capabilities of AI agents in structured data versus their struggles with real-world, messy data like video recordings, sensor telemetry, and robot data.

Dmitry Petrov: Bridging Agents and Physical Data - AI Engineer
Dmitry Petrov: Bridging Agents and Physical Data — from AI Engineer

The Data Disconnect for AI Agents

Petrov opened by citing research from Anthropic and OpenAI, revealing that AI agents, even advanced ones, exhibit low accuracy (around 21%) when dealing with data projects without specific Data Harnesses and context. While these labs have explored multi-layered context approaches for structured business logic, Petrov emphasized his work in the 'extreme side of the data universe' with unstructured, physical data. He noted that his decade of experience, including building DVC (Data Version Control), has led him to DataChain and a focus on creating practical solutions for these challenges.

Building the 'Body' for AI: The Data Harness

Petrov argued that for AI agents to truly interact with the physical world, they need more than just a 'brain' (the LLM). They require a 'body', a data harness that enables them to 'see' data properly, 'run' computations, 'verify' results, and 'remember' crucial information. This harness should allow agents to touch data, run tests, and efficiently process complex binary files that often reside in object storages. He illustrated the sheer complexity of such data, where 2,000 video files can easily generate millions of objects, creating a 'neutron star' of data, small on the surface but immense in mass and information density.

Overcoming Data Spaghetti: JSON vs. Databases

The presentation then delved into common but flawed approaches to managing this complex data. The first, putting metadata in millions of JSON files next to raw data, leads to high latency and inconsistency. The second, using a centralized database, creates a two-system problem with different programming languages and stacks, which is often too complex for researchers. Petrov proposed Pydantic as a solution, allowing the use of a single language (Python) for both data schemas and code, eliminating the 'SQL island' issue and simplifying the transition to structured data through Python transpilers.

DataChain in Action: Analyzing Dashcam Footage

Petrov demonstrated DataChain's open-source project, showcasing how a data harness can be integrated with coding agents. Using a prompt to analyze motion in dashcam recordings, the system guided the user through model selection (YOLO), velocity definition, and granularity settings. The process, which took 24 minutes for 90 videos, resulted in a queryable dataset that could answer questions like, "How many videos have people in it?" The output confirmed that 82 out of 91 clips detected people, highlighting the system's ability to efficiently query and extract insights from the processed data.

The Power of Pydantic and a Unified Stack

The core of DataChain's approach lies in its data models, often Pydantic data classes, which define the structure of the data, including file paths, timestamps, bounding boxes, and nested objects. This structured approach transforms raw data into database rows, enabling rapid analytical queries. Petrov stressed that the problem is specific to unstructured data, unlike the structured data world addressed in many LLM data agent discussions. He also emphasized the importance of an efficient execution engine for crunching terabytes of data, the need for incremental updates and checkpoints to avoid losing work, and the role of running tests for quality control.

Memory and Knowledge Management for Agents

Petrov concluded by discussing the critical aspect of 'memory' for AI agents. He highlighted that while coding agents like Copilot are adept at source code, data agents need a robust context, including dataset descriptions, session context, lineage, and source code, exposed as a knowledge base. This knowledge base, often in Markdown files, allows agents and humans to reuse processed information, preventing redundant computations and wasted resources. DataChain's approach, by connecting raw data, code, and results into a cohesive data lineage, ensures agents 'know everything' about the data, leading to more efficient and intelligent operations.

Ultimately, Petrov's presentation underscored the necessity of building specialized 'data harnesses' that understand the unique 'physics' of physical data, enabling AI agents to overcome the limitations of traditional data processing and unlock the full potential of unstructured data for advanced AI applications.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.