AI Agents at Global Scale Meet Tribal Dungeons

Maersk's Dmitry Buykin says production agents fail in tribal dungeons, and fixing them took 100,000 corrections and a 20:1 SOP corpus.

6 min read
Global shipping operations with AI agents orchestrating state machines
Maersk's agent system runs 200+ instances across country-specific SOPs· AI Engineer
Visual TL;DR
Maersk agent deploymentCore
Dmitry Buykin reports on shipping operations agent built for global scale
From the articleThe Maersk case suggests the moat in production agents is not the model but the SOP corpus and the correction loop that enterprises actually control.
Tribal dungeons emergeDriver
Edge cases where one step fails and legacy systems go incoherent
From the article 2 mentionsThey fail in what Maersk practitioner Dmitry Buykin calls tribal dungeons.
Long tail dominates costsDriver
Exception work by experts across incomplete legacy systems exceeds base workflow
From the articleWhat remains is the long tail where happy paths break.
Loop harness encloses agentCore
System wraps agent with preconditions, back-end calls, validation, and evidence
From the article 2 mentionsSpikes and latencies range from a few minutes to 10 minutes, limited by legacy back ends, not the agent loop.
Corrections beat bigger modelsOutcome
100,000 human corrections outperformed scaling up model parameters for edge cases
From the articleThe Maersk case suggests the moat in production agents is not the model but the SOP corpus and the correction loop that enterprises actually control.
Maersk agent deploymentCore
Dmitry Buykin reports on shipping operations agent built for global scale
From the articleThe Maersk case suggests the moat in production agents is not the model but the SOP corpus and the correction loop that enterprises actually control.
Screenshots fail as SOPsContext
Legacy SOPs are click sequences lacking preconditions, validation, and recovery steps
From the articleBuykin said those legacy SOPs are a bunch of screenshots in sequence.
Tribal dungeons emergeDriver
Edge cases where one step fails and legacy systems go incoherent
From the article 2 mentionsThey fail in what Maersk practitioner Dmitry Buykin calls tribal dungeons.
Long tail dominates costsDriver
Exception work by experts across incomplete legacy systems exceeds base workflow
From the articleWhat remains is the long tail where happy paths break.
Loop harness encloses agentCore
System wraps agent with preconditions, back-end calls, validation, and evidence
From the article 2 mentionsSpikes and latencies range from a few minutes to 10 minutes, limited by legacy back ends, not the agent loop.
Corrections beat bigger modelsOutcome
100,000 human corrections outperformed scaling up model parameters for edge cases
From the articleThe Maersk case suggests the moat in production agents is not the model but the SOP corpus and the correction loop that enterprises actually control.
Harness not freedomContext
Agent constrained inside validated loop rather than given open-ended autonomy
20 to 1 corpus ratioContext
Each corrected workflow requires twenty supporting SOP documents in the knowledge base
From the article 2 mentionsMaersk built three parts: SOP memory organized as a corpus, execution runtime, and feedback capture.
Contents(6)

AI agents at global scale do not fail on happy paths. They fail in what Maersk practitioner Dmitry Buykin calls tribal dungeons.

AI Agents at Global Scale Meet Tribal Dungeons - AI Engineer
AI Agents at Global Scale Meet Tribal Dungeons, from AI Engineer

Buykin detailed the work in a practitioner report for AI Engineer on shipping operations at Tribal Dungeons of Global Shipping: AI Agents at Global Scale. Every shipment is an orchestration of many parallel state machines.

The expensive part is the long tail

The easy majority is already automated in many companies. What remains is the long tail where happy paths break.

One step fails to complete, systems go incoherent, and an expert must orchestrate across multiple incomplete legacy systems. That exception work costs more than the base workflow.

Screenshots are not a process

Standard operating procedures in regulated industries capture what a person sees and clicks. Buykin said those legacy SOPs are a bunch of screenshots in sequence.

An agent SOP needs preconditions, decisions, identifiers, back-end calls, validation, recovery and evidence of successful execution. Experts own the what. Agents own the how.

The signal process depends on many systems being coherent at once. If any step cannot complete, an exception becomes a guardrail.

The system is the loop around the agent

Maersk built three parts: SOP memory organized as a corpus, execution runtime, and feedback capture. The agent loop is not the system. The refining loop around it is.

The corpus is the asset. It is 20 to 1 versus runtime, modified for country conditions, and illustrated by an SAP slide where the same process means different things in different countries.

Production runs over 200 instances concurrently. Spikes and latencies range from a few minutes to 10 minutes, limited by legacy back ends, not the agent loop.

Why corrections beat bigger models

Expert time is the bottleneck, so Maersk built a triage bench that clusters failures. The trace is shared evidence for experts and engineers to agree on what happened.

A correction only counts when it becomes an executable change. That is the line between opinion and production fix.

Quality comes from replaying real examples with disabled write rights and checking if behavior improved. Maersk logged over 100,000 corrections in the last 9 months.

Heat maps turned thousands of traces into priorities. Buykin said turning one red block takes one to two months of effort from engineers and agents together.

Harness for a cage, not freedom

Discovery needs agent freedom. Production needs a cage.

The harness makes dumb mistakes impossible. Wrong workflow triggers classifier eval. Wrong gate triggers right gate. Wrong assumption triggers review.

Buykin said the team does not use MCPs because bloated system responses need distilling. They tune tools through function calling to control quality.

His five moves: make work representable, make execution bounded, make behavior observable, make correction cheap, and make improvement compound. Successful sequences are merged into larger composite tools that roll out across hundreds of countries at once.

Why this matters

This flips the usual agent demo narrative. Model capability is not the constraint. Representing messy operational knowledge safely is.

For enterprises, accuracy was earned one small correction at a time, not designed in one diagram. Pipe coding and spec driven development plateaued early, Buykin showed, and real gains came after.

StartupHub.ai data shows You raised $80M in its 2023 Series A, and the competitors we track in that search and answer space include Perplexity AI, Google (NASDAQ:GOOGL), Lucidworks, matey and Dante. The Maersk case suggests the moat in production agents is not the model but the SOP corpus and the correction loop that enterprises actually control.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.