AI EngineerVideoSource Linked
State of Data — Sean Cai, Independent / State of Data
Real workflow traces may teach agents more than manufactured tasks, while benchmark scores can shift with the harness. Build pipelines around live work and test across scaffolds.
AI Engineer · Jul 26, 2026
Source Summary
Cai separates saved outputs from process data: trajectories, decisions, and reasoning traces. He calls minimally shaped real workflows **Type 1 data** and expert-manufactured examples **Type 2 data**, arguing that realism comes from the work itself.
Practical Implication
For coding agents, capture actual tool use, state transitions, failure recovery, and outcomes. Evaluate across harnesses and infrastructure: a score from **one benchmark under one scaffold** is only one sample, not a stable measure of capability.
Agent-Ready Context
Cai separates saved outputs from process data: trajectories, decisions, and reasoning traces. He calls minimally shaped real workflows **Type 1 data** and expert-manufactured examples **Type 2 data**, arguing that realism comes from the work itself. For coding agents, capture actual tool use, state transitions, failure recovery, and outcomes. Evaluate across harnesses and infrastructure: a score from **one benchmark under one scaffold** is only one sample, not a stable measure of capability. The talk presents a market thesis, not a controlled study. Its claimed **20–30 vendor** diversification and benchmark failure modes are not quantified in the transcript, while domains such as robotics still face unresolved choices about what data modality to collect.
Context Map
benchmarkcodingdata#agent-evals#benchmark-integrity#harness-engineeringUncertainty
The talk presents a market thesis, not a controlled study. Its claimed **20–30 vendor** diversification and benchmark failure modes are not quantified in the transcript, while domains such as robotics still face unresolved choices about what data modality to collect.