Verifiable Environments for AI in Biology — Kenny Workman, LatchBio
Biology agents need evaluators that verify analysis of large experimental datasets, not recall. LatchBio found human review essential because valid scientific paths can defeat brittle graders.
A single-cell run can produce **2–6 TB**, while a spatial biology run can reach **7 TB**. LatchBio’s Spatial Bench contains **146 problems** with data inputs, scientific tasks, grader configuration, and deterministic checks.
Builders of research agents should require conclusions to come from interacting with the supplied data. Ground truth must remain valid across legitimate analysis paths, and human attempts should test whether deterministic graders reject scientifically sound alternatives.
A single-cell run can produce **2–6 TB**, while a spatial biology run can reach **7 TB**. LatchBio’s Spatial Bench contains **146 problems** with data inputs, scientific tasks, grader configuration, and deterministic checks. Builders of research agents should require conclusions to come from interacting with the supplied data. Ground truth must remain valid across legitimate analysis paths, and human attempts should test whether deterministic graders reject scientifically sound alternatives. End-state rewards become weak as workflows grow longer, and current models still miss full biological tasks. Each long-horizon evaluation reportedly took **three people about a week** to create, showing how expensive durable domain verification can be.
This grounds long-horizon scientific-agent evaluation in supplied data, scientifically plural analysis paths, and deterministic checks tested by humans. It reinforces final-state and domain-specific grading, while exposing a hard constraint missing from generic eval recipes: durable ground truth can require week-scale expert labor per task, and deterministic graders can still reject valid science unless alternative workflows are deliberately validated.