# Tribal Dungeons of Global Shipping: AI Agents at Global Scale — Dmitry Buykin, Maersk

Source: [AI Engineer](https://www.youtube.com/watch?v=dQ-_i1tZiws)  
Feed7 permalink: https://feed7.dev/p/tribal-dungeons-of-global-shipping-ai-agents-at-global-scale-dmitry-buyk-0kcigeh  
Published: 2026-08-29T17:30:21.000Z  
Trust: Source Linked (source_linked)

## Why Included

Maersk’s production agents depend less on a clever loop than on executable SOPs, bounded tools, replayable traces, and a correction system shared by experts and engineers.

## Source Summary

Maersk runs **over 200 agent instances** for shipping operations, where legacy systems can stretch latency from minutes to **10 minutes**. Its executable SOP corpus captures preconditions, decisions, calls, validation, recovery, and evidence.

## Practical Implication

Treat the harness and correction loop as the product. Convert screenshots and expert habits into testable procedures, constrain production rights, cluster failures, and replay real cases before promoting changes.

## Agent-Ready Context

Maersk runs **over 200 agent instances** for shipping operations, where legacy systems can stretch latency from minutes to **10 minutes**. Its executable SOP corpus captures preconditions, decisions, calls, validation, recovery, and evidence.

Treat the harness and correction loop as the product. Convert screenshots and expert habits into testable procedures, constrain production rights, cluster failures, and replay real cases before promoting changes.

This approach required **over 100,000 corrections in 9 months**, with expert time still the bottleneck. The talk reports Maersk’s operating method, not a portable benchmark or proof that the same architecture fits smaller workflows.

## Connected Context

Feed7 judgment across 669 accumulated Signals:

This turns familiar harness controls into evidence from a large, latency-heavy production operation: reliability comes from executable procedures, constrained rights, replay, and an institutional correction loop. The volume of corrections and continuing expert bottleneck narrow the automation claim, showing that stronger infrastructure reorganizes domain labor rather than eliminating it, and that the pattern is not automatically portable to smaller workflows.

- [AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTok](https://feed7.dev/p/ai-agents-are-just-distributed-systems-now-salman-munaf-tiktok-1v4yc47) — Maersk’s validation, recovery, evidence, and constrained production rights provide operating evidence for the distributed-systems controls this candidate says are required when agents mutate external state.
- [Twin: Playing an Unknown Game with a Test-Time Digital Twin](https://feed7.dev/p/2608-14490v1-0d3xjvt) — Both make executable replay a promotion gate, but Twin demonstrates it in a cheap simulator while Maersk shows the much heavier correction and expert-maintenance burden of applying replay to production operations.
- [Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute](https://feed7.dev/p/learning-on-the-job-the-future-of-post-training-raymond-feng-applied-com-17u0m7t) — Maersk’s large correction corpus is the kind of production feedback infrastructure this candidate considers valuable, while its slow legacy workflows and expert bottleneck illustrate why such experience is harder to replay and convert into training updates.
- [Your Agent Just Authorized What?! — Jay Mok & Ben Coumes, Paypal](https://feed7.dev/p/your-agent-just-authorized-what-jay-mok-ben-coumes-paypal-024znqi) — Maersk’s constrained production rights reinforce the candidate’s consequence-sensitive authorization principle, adding a concrete operational setting where agent permissions must be bounded and actions leave evidence.

## Context Map

- Layer: agent
- Domains: coding
- Topics: harness-engineering, agent-reliability, tool-use

## Uncertainty

- This approach required **over 100,000 corrections in 9 months**, with expert time still the bottleneck. The talk reports Maersk’s operating method, not a portable benchmark or proof that the same architecture fits smaller workflows.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
