# State of Data — Sean Cai, Independent / State of Data

Source: [AI Engineer](https://www.youtube.com/watch?v=ZyIoTOAbRfs)  
Feed7 permalink: https://feed7.dev/p/state-of-data-sean-cai-independent-state-of-data-0v9fy69  
Published: 2026-07-26T17:00:06.000Z  
Trust: Source Linked (source_linked)

## Why Included

Real workflow traces may teach agents more than manufactured tasks, while benchmark scores can shift with the harness. Build pipelines around live work and test across scaffolds.

## Source Summary

Cai separates saved outputs from process data: trajectories, decisions, and reasoning traces. He calls minimally shaped real workflows **Type 1 data** and expert-manufactured examples **Type 2 data**, arguing that realism comes from the work itself.

## Practical Implication

For coding agents, capture actual tool use, state transitions, failure recovery, and outcomes. Evaluate across harnesses and infrastructure: a score from **one benchmark under one scaffold** is only one sample, not a stable measure of capability.

## Agent-Ready Context

Cai separates saved outputs from process data: trajectories, decisions, and reasoning traces. He calls minimally shaped real workflows **Type 1 data** and expert-manufactured examples **Type 2 data**, arguing that realism comes from the work itself.

For coding agents, capture actual tool use, state transitions, failure recovery, and outcomes. Evaluate across harnesses and infrastructure: a score from **one benchmark under one scaffold** is only one sample, not a stable measure of capability.

The talk presents a market thesis, not a controlled study. Its claimed **20–30 vendor** diversification and benchmark failure modes are not quantified in the transcript, while domains such as robotics still face unresolved choices about what data modality to collect.

## Context Map

- Layer: benchmark
- Domains: coding, data
- Topics: agent-evals, benchmark-integrity, harness-engineering

## Uncertainty

- The talk presents a market thesis, not a controlled study. Its claimed **20–30 vendor** diversification and benchmark failure modes are not quantified in the transcript, while domains such as robotics still face unresolved choices about what data modality to collect.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
