AI Agents at Global Scale Meet Tribal Dungeons
Maersk's Dmitry Buykin says production agents fail in tribal dungeons, and fixing them took 100,000 corrections and a 20:1 SOP corpus.
6 min read

Visual TL;DR
Dmitry Buykin reports on shipping operations agent built for global scale
From the articleThe Maersk case suggests the moat in production agents is not the model but the SOP corpus and the correction loop that enterprises actually control.
Edge cases where one step fails and legacy systems go incoherent
From the article 2 mentionsThey fail in what Maersk practitioner Dmitry Buykin calls tribal dungeons.
Exception work by experts across incomplete legacy systems exceeds base workflow
From the articleWhat remains is the long tail where happy paths break.
System wraps agent with preconditions, back-end calls, validation, and evidence
From the article 2 mentionsSpikes and latencies range from a few minutes to 10 minutes, limited by legacy back ends, not the agent loop.
100,000 human corrections outperformed scaling up model parameters for edge cases
From the articleThe Maersk case suggests the moat in production agents is not the model but the SOP corpus and the correction loop that enterprises actually control.
Dmitry Buykin reports on shipping operations agent built for global scale
From the articleThe Maersk case suggests the moat in production agents is not the model but the SOP corpus and the correction loop that enterprises actually control.
Legacy SOPs are click sequences lacking preconditions, validation, and recovery steps
From the articleBuykin said those legacy SOPs are a bunch of screenshots in sequence.
Edge cases where one step fails and legacy systems go incoherent
From the article 2 mentionsThey fail in what Maersk practitioner Dmitry Buykin calls tribal dungeons.
Exception work by experts across incomplete legacy systems exceeds base workflow
From the articleWhat remains is the long tail where happy paths break.
System wraps agent with preconditions, back-end calls, validation, and evidence
From the article 2 mentionsSpikes and latencies range from a few minutes to 10 minutes, limited by legacy back ends, not the agent loop.
100,000 human corrections outperformed scaling up model parameters for edge cases
From the articleThe Maersk case suggests the moat in production agents is not the model but the SOP corpus and the correction loop that enterprises actually control.
Agent constrained inside validated loop rather than given open-ended autonomy
Each corrected workflow requires twenty supporting SOP documents in the knowledge base
From the article 2 mentionsMaersk built three parts: SOP memory organized as a corpus, execution runtime, and feedback capture.
Contents(6)
© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.