{
  "schema_version": "1.1",
  "id": "archive:https://www.youtube.com/watch?v=Ki980nV0__0",
  "slug": "computer-use-models-will-agentify-the-web-not-apis-dhruv-batra-yutori-07zapmo",
  "url": "https://feed7.dev/p/computer-use-models-will-agentify-the-web-not-apis-dhruv-batra-yutori-07zapmo",
  "title": "Computer-use models will agentify the web, not APIs — Dhruv Batra, Yutori",
  "why_included": "The long-tail web is unlikely to expose clean agent APIs. Browser agents need pixels as ground truth, with code and network access used opportunistically for speed rather than as universal substitutes.",
  "summary": "The argument separates popular services with structured endpoints from a long tail of scanned PDFs, image menus, delayed network data, and state expressed only through rendered UI. It estimates roughly **200 million active websites** and about **1 billion total websites**.",
  "practical_implication": "Build computer-use systems that inspect pixels to verify outcomes, but let them read network responses or generate JavaScript when that is more efficient. For broad research, an orchestrator can dispatch multiple browser navigators in separate sandboxes and collect structured results.",
  "agent_context": "The argument separates popular services with structured endpoints from a long tail of scanned PDFs, image menus, delayed network data, and state expressed only through rendered UI. It estimates roughly **200 million active websites** and about **1 billion total websites**.\n\nBuild computer-use systems that inspect pixels to verify outcomes, but let them read network responses or generate JavaScript when that is more efficient. For broad research, an orchestrator can dispatch multiple browser navigators in separate sandboxes and collect structured results.\n\nThe talk says current accuracy differences against frontier models are within statistical noise, while latency and cost are the stronger case for specialized models. The prediction of sub-penny, sub-100 ms browser tasks is an aspiration, not a demonstrated present capability.",
  "source": {
    "name": "AI Engineer",
    "url": "https://www.youtube.com/watch?v=Ki980nV0__0",
    "published_at": "2026-08-14T14:00:06.000Z"
  },
  "source_class": "video",
  "content_type": "Video",
  "layer": "model",
  "domains": [
    "research"
  ],
  "topics": [
    "computer-use",
    "multi-agent",
    "tool-use"
  ],
  "verification": {
    "status": "source_linked",
    "label": "Source Linked",
    "method": "source_feed",
    "verified_at": null
  },
  "uncertainty": [
    "The talk says current accuracy differences against frontier models are within statistical noise, while latency and cost are the stronger case for specialized models. The prediction of sub-penny, sub-100 ms browser tasks is an aspiration, not a demonstrated present capability."
  ],
  "connected_context": {
    "meaning": "This narrows the API-versus-GUI choice into a hybrid access strategy: use rendered pixels for user-visible truth, but exploit network data or generated code when efficient. It extends computer use toward parallel web research over the unstructured long tail, while making cost and latency—not demonstrated accuracy superiority—the present case for specialized models.",
    "corpus_size": 462,
    "generated_at": "2026-08-16T10:03:55.054Z",
    "connections": [
      {
        "title": "The Dark Arts of Web Automation: Teaching Agents to Use Websites Like Humans — Corey Gallon, Rexmore",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=26RtyAm9y_Q",
        "feed7_url": "https://feed7.dev/p/the-dark-arts-of-web-automation-teaching-agents-to-use-websites-like-hum-12nj0e0",
        "reason": "Concretizes the hybrid strategy by combining programmable browser access with human-like interaction and verification, then compiling stable paths into deterministic code."
      },
      {
        "title": "ChromeDevTools/chrome-devtools-mcp",
        "source_name": "GitHub",
        "source_url": "https://github.com/ChromeDevTools/chrome-devtools-mcp",
        "feed7_url": "https://feed7.dev/p/chrome-devtools-mcp-0ow49x2",
        "reason": "Provides an implementation route for jointly inspecting screenshots, network responses, console state, and browser behavior, while exposing the privacy boundary this access creates."
      },
      {
        "title": "Perception Agents — Antje Barth, Amazon AGI Lab",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=2JX6JYyQG4Y",
        "feed7_url": "https://feed7.dev/p/perception-agents-antje-barth-amazon-agi-lab-1gq6jg4",
        "reason": "Reinforces pixels as the verification surface shared with users, especially when structured state fails to capture the rendered outcome."
      },
      {
        "title": "alibaba/page-agent",
        "source_name": "GitHub",
        "source_url": "https://github.com/alibaba/page-agent",
        "feed7_url": "https://feed7.dev/p/page-agent-1exaacx",
        "reason": "Offers a contrasting DOM-text approach for instrumented pages, which may avoid multimodal browser control but does not address the rendered-only and unstructured web cases motivating this talk."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-14T14:00:06.000Z",
  "modified_at": "2026-08-14T14:00:06.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/computer-use-models-will-agentify-the-web-not-apis-dhruv-batra-yutori-07zapmo",
    "json": "https://feed7.dev/p/computer-use-models-will-agentify-the-web-not-apis-dhruv-batra-yutori-07zapmo.json",
    "markdown": "https://feed7.dev/p/computer-use-models-will-agentify-the-web-not-apis-dhruv-batra-yutori-07zapmo.md"
  }
}