SWE-Gate shows why green tests are an incomplete agent-eval signal: 221 of 644 functionally passing repairs still violated constraints derived from code review.FEED7 SUMMARY
A 52,988-request audit found black-box LLM judges too unstable for preregistered gates. Measure repeatability first, and pilot the judge before fixing evaluation thresholds.FEED7 SUMMARY
AI Gateway can now cap each user's aggregate spend across attributed API keys and app tokens, giving unattended coding agents a hard cost boundary.FEED7 SUMMARY
Vercel found that prose alone produced inconsistent agent-made pages, then paired design.md with fixed CSS primitives and repeatable evals to encode brand judgment.FEED7 SUMMARY
OpenAI has deprecated its standalone skills catalog. Builders should use the Plugins repository and current plugin guide for examples, including skill-only packaging.FEED7 SUMMARY
Knowledge-work agents need code-like infrastructure around tools: centralized context, action records, verification, enforced permissions, and preflight checks for irreversible work.FEED7 SUMMARY
agent#harness-engineering
Sep 3, 2026
What to Ignore This Week
GPT-6 Astra: A new generation of intelligence. The announcement lacks benchmarks, pricing, limits, and API details, leaving no evidence-based reason to change an agent default.
# feed7 Weekly #009 — Context
Use this for: preparing an agent-workflow planning session
- SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents (arXiv, needs review)
- Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints (arXiv, needs review)
- Set per-user budgets on AI Gateway (Vercel, official source)