This pack turns common engineering practices into portable coding-agent workflows for specs, TDD, review and shipping. Its useful idea is to require evidence at each gate, not merely better prompts.FEED7 SUMMARY
A production refactor shows why coding-agent evaluations need acceptance criteria and end-to-end verification: fast output can still be incomplete scaffolding.FEED7 SUMMARY
MIST tests whether models use good context while resisting bad context, exposing agents that appear robust only because they ignore external evidence altogether.FEED7 SUMMARY
Across BFCL v4, models usually handled tools as typed Python calls at least as well as native JSON, suggesting code-based orchestration is worth testing for capable coding agents.FEED7 SUMMARY
Cursor Router learns task complexity and model fit from production behavior, showing why agent routing should include correction signals, cache costs, and per-task performance.FEED7 SUMMARY
Chat SDK can pause a workflow for a verified human decision and resume after seconds or days, without a custom approvals table, action handler, or polling loop.FEED7 SUMMARY
tools#agent-sdks
Aug 6, 2026
What to Ignore This Week
New ways to learn and teach with ChatGPT Work and Codex. The announcement names education plugins but provides no plugin list, capabilities, pricing, availability, or integration details.