Sign InOpen Brain
AI EngineerVideoSource Linked

Verifiable Environments for AI in Biology — Kenny Workman, LatchBio

Biology agents need evaluators that verify analysis of large experimental datasets, not recall. LatchBio found human review essential because valid scientific paths can defeat brittle graders.

AI Engineer · Jul 31, 2026
Open Source Open MarkdownOpen JSON
Source Summary

A single-cell run can produce **2–6 TB**, while a spatial biology run can reach **7 TB**. LatchBio’s Spatial Bench contains **146 problems** with data inputs, scientific tasks, grader configuration, and deterministic checks.

Practical Implication

Builders of research agents should require conclusions to come from interacting with the supplied data. Ground truth must remain valid across legitimate analysis paths, and human attempts should test whether deterministic graders reject scientifically sound alternatives.

Agent-Ready Context
A single-cell run can produce **2–6 TB**, while a spatial biology run can reach **7 TB**. LatchBio’s Spatial Bench contains **146 problems** with data inputs, scientific tasks, grader configuration, and deterministic checks.

Builders of research agents should require conclusions to come from interacting with the supplied data. Ground truth must remain valid across legitimate analysis paths, and human attempts should test whether deterministic graders reject scientifically sound alternatives.

End-state rewards become weak as workflows grow longer, and current models still miss full biological tasks. Each long-horizon evaluation reportedly took **three people about a week** to create, showing how expensive durable domain verification can be.
Connected Context · Feed7 Judgment

This grounds long-horizon scientific-agent evaluation in supplied data, scientifically plural analysis paths, and deterministic checks tested by humans. It reinforces final-state and domain-specific grading, while exposing a hard constraint missing from generic eval recipes: durable ground truth can require week-scale expert labor per task, and deterministic graders can still reject valid science unless alternative workflows are deliberately validated.

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta SoftwareBoth reject duration alone as an adequate long-horizon measure and favor inspectable environments and outcomes; this target adds the need to preserve multiple scientifically valid paths through deterministic grading.onepot-Bench 0: towards lab-aware in silico chemistry benchmarksBoth specialize benchmark integrity around lab-grounded scientific work; Spatial Bench adds explicit supplied-data interaction and human testing of whether deterministic checks reject legitimate analyses.Demystifying evals for AI agentsThe start-small eval roadmap is compatible with the high expert cost reported here, but this target shows why scaling beyond a small task set can remain labor-intensive in scientific domains.Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge BiasThe candidate’s evidence of LLM-judge bias supports the target’s preference for deterministic checks, while the target shows that deterministic grading has its own failure mode when valid scientific alternatives are not encoded.
Context Map
benchmarkresearchdata#agent-evals#benchmark-integrity#agent-reliability
Uncertainty
End-state rewards become weak as workflows grow longer, and current models still miss full biological tasks. Each long-horizon evaluation reportedly took **three people about a week** to create, showing how expensive durable domain verification can be.