{
  "schema_version": "1.1",
  "id": "weekly-2026-08-17",
  "slug": "2026-08-17",
  "issue_number": "006",
  "title": "feed7 Weekly #006",
  "description": "Six practical upgrades for faster, safer, and more measurable coding-agent sessions.",
  "published_at": "2026-08-17T00:00:00.000Z",
  "modified_at": "2026-08-16T18:04:36.048Z",
  "url": "https://feed7.dev/weekly/2026-08-17",
  "formats": {
    "html": "https://feed7.dev/weekly/2026-08-17",
    "json": "https://feed7.dev/weekly/2026-08-17.json",
    "markdown": "https://feed7.dev/weekly/2026-08-17.md"
  },
  "selection": {
    "rule": "Six source-backed signals and one distraction to leave out.",
    "mode": "ai",
    "ignore_item_id": "auto-fb21820d80"
  },
  "items": [
    {
      "schema_version": "1.1",
      "id": "auto-31ceb9467c",
      "slug": "quotebench-how-matched-scores-can-hide-command-path-fail-31ceb9467c",
      "url": "https://feed7.dev/p/quotebench-how-matched-scores-can-hide-command-path-fail-31ceb9467c",
      "title": "QuoteBench: How Matched Scores Can Hide Command-Path Failures",
      "why_included": "Test commands through the exact production transport and verify final state, since one parser cut task completion by 55.4–73.2 points.",
      "summary": "QuoteBench shows that shell-command scores can conceal failures introduced by serialization and reparsing. Agent evals should identify the execution path, not attribute every result to the model.",
      "practical_implication": "When the boundary was disclosed, six configurations recovered 30.4–60.7 points. Builders should test generated commands through the exact production transport and validate final state, especially where wrappers interpolate or reparse shell text.",
      "agent_context": "QuoteBench tests **56 one-shot tasks** from 14 incident-derived families across eight configurations. Adding one unescaped parser cut replayed-command success by **55.4–73.2 percentage points**.\n\nWhen the boundary was disclosed, six configurations recovered **30.4–60.7 points**. Builders should test generated commands through the exact production transport and validate final state, especially where wrappers interpolate or reparse shell text.\n\nAdaptation was uneven: two configurations recovered nothing or declined slightly. GPT-5.6-sol’s **−3.6-point matched gap** concealed −64.3 points of transport damage plus 60.7 points of compensation, so aggregate scores can misstate both model and harness quality.",
      "source": {
        "name": "arXiv",
        "url": "https://arxiv.org/abs/2608.13547v1",
        "published_at": "2026-08-13T00:00:00.000Z"
      },
      "source_class": "blog_post",
      "content_type": "Paper",
      "layer": "benchmark",
      "domains": [
        "coding",
        "security"
      ],
      "topics": [
        "agent-evals",
        "benchmark-integrity",
        "harness-engineering"
      ],
      "verification": {
        "status": "needs_review",
        "label": "Needs Review",
        "method": "unverified",
        "verified_at": null
      },
      "uncertainty": [
        "Automatically selected from source material; feed7 has not independently tested the claim."
      ],
      "connected_context": null,
      "lifecycle": "New",
      "published_at": "2026-08-13T00:00:00.000Z",
      "modified_at": "2026-08-13T00:00:00.000Z",
      "supersedes": [],
      "expires_at": null,
      "formats": {
        "html": "https://feed7.dev/p/quotebench-how-matched-scores-can-hide-command-path-fail-31ceb9467c",
        "json": "https://feed7.dev/p/quotebench-how-matched-scores-can-hide-command-path-fail-31ceb9467c.json",
        "markdown": "https://feed7.dev/p/quotebench-how-matched-scores-can-hide-command-path-fail-31ceb9467c.md"
      }
    },
    {
      "schema_version": "1.1",
      "id": "auto-a24f8f74a4",
      "slug": "computer-use-at-the-edge-of-the-statistical-precipice-pi-a24f8f74a4",
      "url": "https://feed7.dev/p/computer-use-at-the-edge-of-the-statistical-precipice-pi-a24f8f74a4",
      "title": "Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs",
      "why_included": "Vary task data, appearance, and initial state so agent evals measure adaptation instead of rewarding replayed action scripts.",
      "summary": "Static computer-use benchmarks can reward memorized action scripts rather than adaptation. Vary task state, verify every generated case, and calculate uncertainty across both actions and environments.",
      "practical_implication": "For agent evals, vary data, appearance, and initial state; automatically reject invalid combinations; and use privileged verifiers inside a sandbox. DGWorld applies this design across 15 apps, 387 scenarios, and 3.2 million verified configurations.",
      "agent_context": "A replay agent stores one winning trajectory per task and blindly repeats it. On deterministic OSWorld and MobileWorld-style evaluations, a script **under 1 MB** can match or beat the frontier model that generated its traces, exposing benchmark replayability.\n\nFor agent evals, vary data, appearance, and initial state; automatically reject invalid combinations; and use privileged verifiers inside a sandbox. DGWorld applies this design across **15 apps, 387 scenarios, and 3.2 million verified configurations**.\n\nUncertainty must cover both model actions and environment variation. The talk reports that rollout-only intervals can provide roughly **17–20% coverage** where a correctly structured method approaches the intended 95%, but applying that method requires a benchmark with explicit variation structure.",
      "source": {
        "name": "AI Engineer",
        "url": "https://www.youtube.com/watch?v=CTLa_p6iOiY",
        "published_at": "2026-08-14T00:00:00.000Z"
      },
      "source_class": "video",
      "content_type": "Video",
      "layer": "benchmark",
      "domains": [],
      "topics": [
        "agent-evals",
        "benchmark-integrity",
        "agent-reliability"
      ],
      "verification": {
        "status": "source_linked",
        "label": "Source Linked",
        "method": "source_feed",
        "verified_at": null
      },
      "uncertainty": [],
      "connected_context": null,
      "lifecycle": "New",
      "published_at": "2026-08-14T00:00:00.000Z",
      "modified_at": "2026-08-14T00:00:00.000Z",
      "supersedes": [],
      "expires_at": null,
      "formats": {
        "html": "https://feed7.dev/p/computer-use-at-the-edge-of-the-statistical-precipice-pi-a24f8f74a4",
        "json": "https://feed7.dev/p/computer-use-at-the-edge-of-the-statistical-precipice-pi-a24f8f74a4.json",
        "markdown": "https://feed7.dev/p/computer-use-at-the-edge-of-the-statistical-precipice-pi-a24f8f74a4.md"
      }
    },
    {
      "schema_version": "1.1",
      "id": "auto-73fc37ffd6",
      "slug": "building-a-software-factory-for-ai-sdk-73fc37ffd6",
      "url": "https://feed7.dev/p/building-a-software-factory-for-ai-sdk-73fc37ffd6",
      "title": "Building a software factory for AI SDK",
      "why_included": "Give each agent one reviewable task with isolated context and evidence, while keeping human scrutiny proportional to change risk.",
      "summary": "Vercel’s AI SDK factory shows a practical scaling pattern: narrow agents produce evidence inside sandboxes while humans retain merge authority and review effort follows risk.",
      "practical_implication": "The reusable pattern is one agent per reviewable task, each with its own prompt, context, and evals. Start locally, pass evidence between classification, analysis, reproduction, implementation, and review stages, then vary human scrutiny by change risk rather than treating every agent output equally.",
      "agent_context": "AI SDK had accumulated **over 1,000 issues and almost 800 pull requests** by late June. Four weeks after introducing its factory, Vercel says agents authored **25–35% of merged PRs** and closed **70–80% of issues**, while humans approved every merge.\n\nThe reusable pattern is one agent per reviewable task, each with its own prompt, context, and evals. Start locally, pass evidence between classification, analysis, reproduction, implementation, and review stages, then vary human scrutiny by change risk rather than treating every agent output equally.\n\nThese are early results from one large open-source project, not a controlled comparison. The factory also depends on isolated sandboxes, restricted secrets and networking, queues, monitoring, and sustained human review, so the headline automation rates omit substantial operating machinery.",
      "source": {
        "name": "Vercel",
        "url": "https://vercel.com/blog/building-a-software-factory-for-ai-sdk",
        "published_at": "2026-08-12T00:00:00.000Z"
      },
      "source_class": "blog_post",
      "content_type": "Engineering Post",
      "layer": "agent",
      "domains": [
        "coding",
        "security"
      ],
      "topics": [
        "multi-agent",
        "harness-engineering",
        "sandboxing"
      ],
      "verification": {
        "status": "official_source",
        "label": "Official Source",
        "method": "source_feed",
        "verified_at": null
      },
      "uncertainty": [],
      "connected_context": null,
      "lifecycle": "New",
      "published_at": "2026-08-12T00:00:00.000Z",
      "modified_at": "2026-08-12T00:00:00.000Z",
      "supersedes": [],
      "expires_at": null,
      "formats": {
        "html": "https://feed7.dev/p/building-a-software-factory-for-ai-sdk-73fc37ffd6",
        "json": "https://feed7.dev/p/building-a-software-factory-for-ai-sdk-73fc37ffd6.json",
        "markdown": "https://feed7.dev/p/building-a-software-factory-for-ai-sdk-73fc37ffd6.md"
      }
    },
    {
      "schema_version": "1.1",
      "id": "auto-fec4301fea",
      "slug": "cloud-agents-start-3x-faster-with-builds-fec4301fea",
      "url": "https://feed7.dev/p/cloud-agents-start-3x-faster-with-builds-fec4301fea",
      "title": "Cloud agents start 3x faster with builds",
      "why_included": "Prebuild deterministic dependencies, start session-fresh services separately, and trace each run to its build and commit SHA.",
      "summary": "Cursor Cloud Agents can start from continuously prepared environment snapshots instead of reinstalling each session. Internal time to first token improved 3x, with failed builds falling back to the last good state.",
      "practical_implication": "Builders should move deterministic setup into the install command, keep session-fresh services in the start command, and use team or environment secrets for private registries. Agent runs can be traced to exact builds and commit SHAs.",
      "agent_context": "Cursor’s new **builds** continuously prepare cloud-agent environments with repositories, dependencies, and install scripts already completed. Internally, environments booted **10x faster** and time to first token improved **3x**.\n\nBuilders should move deterministic setup into the install command, keep session-fresh services in the start command, and use team or environment secrets for private registries. Agent runs can be traced to exact builds and commit SHAs.\n\nThe gains are Cursor’s internal measurements, and user results will depend on repository setup. Builds run hourly by default; failed builds are rejected, leaving agents on the last working snapshot, which may be older than the default branch.",
      "source": {
        "name": "Cursor",
        "url": "https://cursor.com/blog/builds",
        "published_at": "2026-08-13T00:00:00.000Z"
      },
      "source_class": "blog_post",
      "content_type": "Engineering Post",
      "layer": "infra",
      "domains": [
        "coding"
      ],
      "topics": [
        "cloud-agents",
        "sandboxing",
        "observability"
      ],
      "verification": {
        "status": "official_source",
        "label": "Official Source",
        "method": "source_feed",
        "verified_at": null
      },
      "uncertainty": [],
      "connected_context": null,
      "lifecycle": "New",
      "published_at": "2026-08-13T00:00:00.000Z",
      "modified_at": "2026-08-13T00:00:00.000Z",
      "supersedes": [],
      "expires_at": null,
      "formats": {
        "html": "https://feed7.dev/p/cloud-agents-start-3x-faster-with-builds-fec4301fea",
        "json": "https://feed7.dev/p/cloud-agents-start-3x-faster-with-builds-fec4301fea.json",
        "markdown": "https://feed7.dev/p/cloud-agents-start-3x-faster-with-builds-fec4301fea.md"
      }
    },
    {
      "schema_version": "1.1",
      "id": "auto-e7ee7f4465",
      "slug": "lessons-from-studying-every-memory-system-shlok-khemani--e7ee7f4465",
      "url": "https://feed7.dev/p/lessons-from-studying-every-memory-system-shlok-khemani--e7ee7f4465",
      "title": "Lessons from Studying Every Memory System — Shlok Khemani, Independent",
      "why_included": "Make agent memory inspectable and editable, with conflict detection and an explicit update cadence to prevent false beliefs from persisting.",
      "summary": "Persistent memory is a compute and product tradeoff, not just retrieval. Profiles need conflict detection, visibility, editing, and deliberate update cadence before agents can rely on them.",
      "practical_implication": "For agent memory, choose update frequency and profile size as an explicit compute budget. Make stored beliefs inspectable and editable, preserve source context where possible, and detect uncertainty or contradictions before a profile silently steers future work.",
      "agent_context": "ChatGPT and Claude have converged on running user profiles plus tools for retrieving past conversations. The observed implementations differ: ChatGPT's profile is about **4,000 tokens** and updates every few days, while Claude's is about **1,000 tokens** and updates every 24 hours.\n\nFor agent memory, choose update frequency and profile size as an explicit compute budget. Make stored beliefs inspectable and editable, preserve source context where possible, and detect uncertainty or contradictions before a profile silently steers future work.\n\nThe Turkey example shows the core failure: conversations about possible travel were condensed into a trip that never happened. Neither profile summaries nor retrieval alone solve missing evidence, cross-product context silos, or reconciliation with email, calendars, and other sources.",
      "source": {
        "name": "AI Engineer",
        "url": "https://www.youtube.com/watch?v=5ZGyKWjQDr0",
        "published_at": "2026-08-12T00:00:00.000Z"
      },
      "source_class": "video",
      "content_type": "Video",
      "layer": "agent",
      "domains": [],
      "topics": [
        "agent-memory",
        "context-engineering",
        "retrieval"
      ],
      "verification": {
        "status": "source_linked",
        "label": "Source Linked",
        "method": "source_feed",
        "verified_at": null
      },
      "uncertainty": [],
      "connected_context": null,
      "lifecycle": "New",
      "published_at": "2026-08-12T00:00:00.000Z",
      "modified_at": "2026-08-12T00:00:00.000Z",
      "supersedes": [],
      "expires_at": null,
      "formats": {
        "html": "https://feed7.dev/p/lessons-from-studying-every-memory-system-shlok-khemani--e7ee7f4465",
        "json": "https://feed7.dev/p/lessons-from-studying-every-memory-system-shlok-khemani--e7ee7f4465.json",
        "markdown": "https://feed7.dev/p/lessons-from-studying-every-memory-system-shlok-khemani--e7ee7f4465.md"
      }
    },
    {
      "schema_version": "1.1",
      "id": "auto-a44214149b",
      "slug": "use-acp-compatible-harnesses-with-the-ai-sdk-harness-lay-a44214149b",
      "url": "https://feed7.dev/p/use-acp-compatible-harnesses-with-the-ai-sdk-harness-lay-a44214149b",
      "title": "Use ACP-compatible harnesses with the AI SDK harness layer",
      "why_included": "Use ACP for portable harness integration, but test permissions, events, and runtime-specific controls before replacing a direct adapter.",
      "summary": "A new ACP meta-adapter lets HarnessAgent run compatible runtimes that lack dedicated integrations. Keep direct Claude Code and Codex adapters where tighter behavior matters.",
      "practical_implication": "Use the meta-adapter to connect a compatible harness that has no native AI SDK integration. For Claude Code and Codex, Vercel recommends their dedicated adapters because they can expose runtime behavior more precisely.",
      "agent_context": "The new **@ai-sdk/harness-acp** package adapts the Agent Client Protocol rather than one runtime. Builders pass an ACP-compatible package to **createACP**, map the harness basics, and use the result with HarnessAgent.\n\nUse the meta-adapter to connect a compatible harness that has no native AI SDK integration. For **Claude Code and Codex**, Vercel recommends their dedicated adapters because they can expose runtime behavior more precisely.\n\nACP is not universal, and its abstraction may limit or alter access to a harness’s internals. Treat protocol compatibility as a portability option, then test permissions, events, and runtime-specific behavior before replacing a direct adapter.",
      "source": {
        "name": "Vercel",
        "url": "https://vercel.com/changelog/use-acp-compatible-harnesses-with-the-ai-sdk-harness-layer",
        "published_at": "2026-08-13T00:00:00.000Z"
      },
      "source_class": "blog_post",
      "content_type": "Engineering Post",
      "layer": "tools",
      "domains": [
        "coding"
      ],
      "topics": [
        "coding-agents",
        "agent-sdks",
        "harness-engineering"
      ],
      "verification": {
        "status": "official_source",
        "label": "Official Source",
        "method": "source_feed",
        "verified_at": null
      },
      "uncertainty": [],
      "connected_context": null,
      "lifecycle": "New",
      "published_at": "2026-08-13T00:00:00.000Z",
      "modified_at": "2026-08-13T00:00:00.000Z",
      "supersedes": [],
      "expires_at": null,
      "formats": {
        "html": "https://feed7.dev/p/use-acp-compatible-harnesses-with-the-ai-sdk-harness-lay-a44214149b",
        "json": "https://feed7.dev/p/use-acp-compatible-harnesses-with-the-ai-sdk-harness-lay-a44214149b.json",
        "markdown": "https://feed7.dev/p/use-acp-compatible-harnesses-with-the-ai-sdk-harness-lay-a44214149b.md"
      }
    },
    {
      "schema_version": "1.1",
      "id": "auto-fb21820d80",
      "slug": "from-assistance-to-execution-how-enterprises-put-ai-to-w-fb21820d80",
      "url": "https://feed7.dev/p/from-assistance-to-execution-how-enterprises-put-ai-to-w-fb21820d80",
      "title": "From assistance to execution: How enterprises put AI to work",
      "why_included": "The adoption narrative lacks sample size, methodology, workflow examples, and implementation guidance for improving an agent session.",
      "summary": "OpenAI’s research frames enterprise AI adoption as a shift from assistance toward agent execution, with ChatGPT and Codex as examples rather than implementation guidance.",
      "practical_implication": "Builders should treat this as evidence of organizational demand for agents that execute work, then look for the underlying research before changing product or deployment plans.",
      "agent_context": "OpenAI reports that enterprises are adopting **agentic AI** through products including **ChatGPT and Codex**, and says a group it calls frontier firms is moving ahead in adoption.\n\nBuilders should treat this as evidence of organizational demand for agents that execute work, then look for the underlying research before changing product or deployment plans.\n\nThe supplied material provides no sample size, methodology, adoption rates, workflow examples, or definition of a frontier firm, so the size of the claimed gap is unclear.",
      "source": {
        "name": "OpenAI",
        "url": "https://openai.com/index/how-enterprises-put-ai-to-work",
        "published_at": "2026-08-12T00:00:00.000Z"
      },
      "source_class": "blog_post",
      "content_type": "Official Release",
      "layer": "industry",
      "domains": [
        "coding"
      ],
      "topics": [
        "adoption",
        "enterprise"
      ],
      "verification": {
        "status": "official_source",
        "label": "Official Source",
        "method": "source_feed",
        "verified_at": null
      },
      "uncertainty": [],
      "connected_context": null,
      "lifecycle": "New",
      "published_at": "2026-08-12T00:00:00.000Z",
      "modified_at": "2026-08-12T00:00:00.000Z",
      "supersedes": [],
      "expires_at": null,
      "formats": {
        "html": "https://feed7.dev/p/from-assistance-to-execution-how-enterprises-put-ai-to-w-fb21820d80",
        "json": "https://feed7.dev/p/from-assistance-to-execution-how-enterprises-put-ai-to-w-fb21820d80.json",
        "markdown": "https://feed7.dev/p/from-assistance-to-execution-how-enterprises-put-ai-to-w-fb21820d80.md"
      }
    }
  ],
  "agent_instruction": "Use these items as source-backed context. Do not invent claims beyond linked material. Prefer practical implications for solo developer work. If sources conflict, call it out."
}