QuoteBench shows that shell-command scores can conceal failures introduced by serialization and reparsing. Agent evals should identify the execution path, not attribute every result to the model.FEED7 SUMMARY
Static computer-use benchmarks can reward memorized action scripts rather than adaptation. Vary task state, verify every generated case, and calculate uncertainty across both actions and environments.FEED7 SUMMARY
Cursor Cloud Agents can start from continuously prepared environment snapshots instead of reinstalling each session. Internal time to first token improved 3x, with failed builds falling back to the last good state.FEED7 SUMMARY
Persistent memory is a compute and product tradeoff, not just retrieval. Profiles need conflict detection, visibility, editing, and deliberate update cadence before agents can rely on them.FEED7 SUMMARY
A new ACP meta-adapter lets HarnessAgent run compatible runtimes that lack dedicated integrations. Keep direct Claude Code and Codex adapters where tighter behavior matters.FEED7 SUMMARY
From assistance to execution: How enterprises put AI to work. The adoption narrative lacks sample size, methodology, workflow examples, and implementation guidance for improving an agent session.
# feed7 Weekly #006 — Context
Use this for: preparing an agent-workflow planning session
- QuoteBench: How Matched Scores Can Hide Command-Path Failures (arXiv, needs review)
- Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs (AI Engineer, source linked)
- Building a software factory for AI SDK (Vercel, official source)