State of Data — Sean Cai, Independent / State of Data
Real workflow traces may teach agents more than manufactured tasks, while benchmark scores can shift with the harness. Build pipelines around live work and test across scaffolds.
Cai separates saved outputs from process data: trajectories, decisions, and reasoning traces. He calls minimally shaped real workflows **Type 1 data** and expert-manufactured examples **Type 2 data**, arguing that realism comes from the work itself.
For coding agents, capture actual tool use, state transitions, failure recovery, and outcomes. Evaluate across harnesses and infrastructure: a score from **one benchmark under one scaffold** is only one sample, not a stable measure of capability.
Cai separates saved outputs from process data: trajectories, decisions, and reasoning traces. He calls minimally shaped real workflows **Type 1 data** and expert-manufactured examples **Type 2 data**, arguing that realism comes from the work itself. For coding agents, capture actual tool use, state transitions, failure recovery, and outcomes. Evaluate across harnesses and infrastructure: a score from **one benchmark under one scaffold** is only one sample, not a stable measure of capability. The talk presents a market thesis, not a controlled study. Its claimed **20–30 vendor** diversification and benchmark failure modes are not quantified in the transcript, while domains such as robotics still face unresolved choices about what data modality to collect.
This shifts evaluation emphasis from saved answers and isolated benchmarks toward real trajectories: tool calls, state changes, recovery, and outcomes collected from live work. It narrows any single score to a scaffold-dependent observation and favors cross-harness testing, but remains a market thesis rather than quantified evidence, leaving modality and collection choices unresolved in some domains.