Sign InOpen Brain
AI EngineerVideoSource Linked

State of Data — Sean Cai, Independent / State of Data

Real workflow traces may teach agents more than manufactured tasks, while benchmark scores can shift with the harness. Build pipelines around live work and test across scaffolds.

AI Engineer · Jul 26, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Cai separates saved outputs from process data: trajectories, decisions, and reasoning traces. He calls minimally shaped real workflows **Type 1 data** and expert-manufactured examples **Type 2 data**, arguing that realism comes from the work itself.

Practical Implication

For coding agents, capture actual tool use, state transitions, failure recovery, and outcomes. Evaluate across harnesses and infrastructure: a score from **one benchmark under one scaffold** is only one sample, not a stable measure of capability.

Agent-Ready Context
Cai separates saved outputs from process data: trajectories, decisions, and reasoning traces. He calls minimally shaped real workflows **Type 1 data** and expert-manufactured examples **Type 2 data**, arguing that realism comes from the work itself.

For coding agents, capture actual tool use, state transitions, failure recovery, and outcomes. Evaluate across harnesses and infrastructure: a score from **one benchmark under one scaffold** is only one sample, not a stable measure of capability.

The talk presents a market thesis, not a controlled study. Its claimed **20–30 vendor** diversification and benchmark failure modes are not quantified in the transcript, while domains such as robotics still face unresolved choices about what data modality to collect.
Connected Context · Feed7 Judgment

This shifts evaluation emphasis from saved answers and isolated benchmarks toward real trajectories: tool calls, state changes, recovery, and outcomes collected from live work. It narrows any single score to a scaffold-dependent observation and favors cross-harness testing, but remains a market thesis rather than quantified evidence, leaving modality and collection choices unresolved in some domains.

DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, DatacurveDeepSWE partially implements the thesis through original sustained-repository tasks and trace inspection, while its missing routine task categories show why one improved benchmark remains insufficient.From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AITrace-to-simulation reconstruction turns Cai’s call to collect real process data into replayable, fixed-condition evaluations with operational release metrics.Quantifying infrastructure noise in agentic coding evalsThe measured score swing across resource configurations provides empirical support for Cai’s claim that infrastructure and scaffolding can materially alter an observed capability score.The Unreasonable Effectiveness of Separating the Task from the Model — Maxime Rivest & Isaac MillerSeparating a task contract from models, prompts, tools, and harnesses provides an implementation structure for the cross-harness comparisons Cai recommends.
Context Map
benchmarkcodingdata#agent-evals#benchmark-integrity#harness-engineering
Uncertainty
The talk presents a market thesis, not a controlled study. Its claimed **20–30 vendor** diversification and benchmark failure modes are not quantified in the transcript, while domains such as robotics still face unresolved choices about what data modality to collect.