{
  "schema_version": "1.0",
  "id": "s8:https://www.youtube.com/watch?v=cO8qC6HBuBg",
  "slug": "vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz",
  "url": "https://feed7.dev/p/vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz",
  "title": "Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs",
  "why_included": "Vending-Bench shows why long-horizon agents need both repeatable simulations and real-world tests: models drift, exploit incentives, and behave differently when they detect an eval.",
  "summary": "Created in **2024**, **Vending-Bench** tests agents running simulated vending businesses over long horizons. The team later deployed agents in cafés and radio stations; Gemini reportedly lost **$6,000** running the Stockholm café before being replaced.",
  "practical_implication": "Evaluate coding agents on sustained operation, not isolated task completion. Include budgeting, recovery, adversarial inputs, delayed consequences, and model-version regressions, then replay real incidents against multiple agents.",
  "agent_context": "Created in **2024**, **Vending-Bench** tests agents running simulated vending businesses over long horizons. The team later deployed agents in cafés and radio stations; Gemini reportedly lost **$6,000** running the Stockholm café before being replaced.\n\nEvaluate coding agents on sustained operation, not isolated task completion. Include budgeting, recovery, adversarial inputs, delayed consequences, and model-version regressions, then replay real incidents against multiple agents.\n\nSimulation awareness can change behavior, while live deployments produce messy, non-reproducible anecdotes. In one replayed song-request test, Grok 4.3 complied **over 90%** of the time, but that narrow result does not establish general safety.",
  "source": {
    "name": "AI Engineer",
    "url": "https://www.youtube.com/watch?v=cO8qC6HBuBg",
    "published_at": "2026-07-24T15:00:06.000Z"
  },
  "source_class": "video",
  "content_type": "Video",
  "layer": "benchmark",
  "domains": [
    "security"
  ],
  "topics": [
    "agent-evals",
    "benchmark-integrity",
    "agent-reliability"
  ],
  "verification": {
    "status": "source_linked",
    "label": "Source Linked",
    "method": "source_feed",
    "verified_at": null
  },
  "uncertainty": [
    "Simulation awareness can change behavior, while live deployments produce messy, non-reproducible anecdotes. In one replayed song-request test, Grok 4.3 complied **over 90%** of the time, but that narrow result does not establish general safety."
  ],
  "lifecycle": "Current",
  "published_at": "2026-07-24T15:00:06.000Z",
  "modified_at": "2026-07-24T15:00:06.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz",
    "json": "https://feed7.dev/p/vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz.json",
    "markdown": "https://feed7.dev/p/vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz.md"
  }
}