{
  "schema_version": "1.1",
  "id": "archive:https://www.youtube.com/watch?v=CTLa_p6iOiY",
  "slug": "computer-use-at-the-edge-of-the-statistical-precipice-pierluca-d-oro-pro-16136hs",
  "url": "https://feed7.dev/p/computer-use-at-the-edge-of-the-statistical-precipice-pierluca-d-oro-pro-16136hs",
  "title": "Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs",
  "why_included": "Static computer-use benchmarks can reward memorized action scripts rather than adaptation. Vary task state, verify every generated case, and calculate uncertainty across both actions and environments.",
  "summary": "A replay agent stores one winning trajectory per task and blindly repeats it. On deterministic OSWorld and MobileWorld-style evaluations, a script **under 1 MB** can match or beat the frontier model that generated its traces, exposing benchmark replayability.",
  "practical_implication": "For agent evals, vary data, appearance, and initial state; automatically reject invalid combinations; and use privileged verifiers inside a sandbox. DGWorld applies this design across **15 apps, 387 scenarios, and 3.2 million verified configurations**.",
  "agent_context": "A replay agent stores one winning trajectory per task and blindly repeats it. On deterministic OSWorld and MobileWorld-style evaluations, a script **under 1 MB** can match or beat the frontier model that generated its traces, exposing benchmark replayability.\n\nFor agent evals, vary data, appearance, and initial state; automatically reject invalid combinations; and use privileged verifiers inside a sandbox. DGWorld applies this design across **15 apps, 387 scenarios, and 3.2 million verified configurations**.\n\nUncertainty must cover both model actions and environment variation. The talk reports that rollout-only intervals can provide roughly **17–20% coverage** where a correctly structured method approaches the intended 95%, but applying that method requires a benchmark with explicit variation structure.",
  "source": {
    "name": "AI Engineer",
    "url": "https://www.youtube.com/watch?v=CTLa_p6iOiY",
    "published_at": "2026-08-14T14:30:31.000Z"
  },
  "source_class": "video",
  "content_type": "Video",
  "layer": "benchmark",
  "domains": [],
  "topics": [
    "agent-evals",
    "benchmark-integrity",
    "agent-reliability"
  ],
  "verification": {
    "status": "source_linked",
    "label": "Source Linked",
    "method": "source_feed",
    "verified_at": null
  },
  "uncertainty": [
    "Uncertainty must cover both model actions and environment variation. The talk reports that rollout-only intervals can provide roughly **17–20% coverage** where a correctly structured method approaches the intended 95%, but applying that method requires a benchmark with explicit variation structure."
  ],
  "connected_context": {
    "meaning": "This demonstrates that fixed computer-use benchmarks can measure trajectory memorization rather than general capability. It strengthens the case for varied initial states, privileged final-state verification, and uncertainty estimates that include environment variation, while limiting the proposed statistical remedy to benchmarks whose variation is explicitly structured.",
    "corpus_size": 462,
    "generated_at": "2026-08-16T10:03:55.054Z",
    "connections": [
      {
        "title": "When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=-npY6XjM8CQ",
        "feed7_url": "https://feed7.dev/p/when-will-the-benchmaxxing-plague-end-nick-heiner-surge-ai-178gqcg",
        "reason": "Supplies a concrete replay mechanism behind the broader warning that leaderboard gains can arise from contamination, weak test conditions, and shortcuts rather than useful behavior."
      },
      {
        "title": "Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=2aS7aKoXn64",
        "feed7_url": "https://feed7.dev/p/rethinking-environments-for-long-horizon-work-rayan-garg-theta-software-11r7wbx",
        "reason": "Reinforces final-state verification but adds that difficulty and uncertainty must be measured across changing environments, not inferred from one fixed trajectory or task duration."
      },
      {
        "title": "Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=cO8qC6HBuBg",
        "feed7_url": "https://feed7.dev/p/vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz",
        "reason": "Supports the need for repeatable yet varied simulations because agents can exploit fixed incentives and behave differently when they recognize evaluation conditions."
      },
      {
        "title": "Verifiable Environments for AI in Biology — Kenny Workman, LatchBio",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=3ZMUiFaQ3qg",
        "feed7_url": "https://feed7.dev/p/verifiable-environments-for-ai-in-biology-kenny-workman-latchbio-1vs6y66",
        "reason": "Qualifies privileged deterministic verification: even strong internal graders can reject valid solution paths unless humans test the acceptable variation."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-14T14:30:31.000Z",
  "modified_at": "2026-08-14T14:30:31.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/computer-use-at-the-edge-of-the-statistical-precipice-pierluca-d-oro-pro-16136hs",
    "json": "https://feed7.dev/p/computer-use-at-the-edge-of-the-statistical-precipice-pierluca-d-oro-pro-16136hs.json",
    "markdown": "https://feed7.dev/p/computer-use-at-the-edge-of-the-statistical-precipice-pierluca-d-oro-pro-16136hs.md"
  }
}