Back
arXivPaperNeeds ReviewbenchmarkNew
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
Before using an LLM judge as a gate, replay identical inputs, measure its noise floor, and pilot thresholds on a small sample.
arXivSep 3, 20262 min
Source Summary
A 52,988-request audit found black-box LLM judges too unstable for preregistered gates. Measure repeatability first, and pilot the judge before fixing evaluation thresholds.
Practical Implication
If an LLM judge gates agent releases or training data, test it as a measurement instrument first. Repeat identical requests, estimate the noise floor, retain execution records, and run a small pilot; the authors say roughly 2% of call volume would have revealed both unreachable gates.
Agent-Ready Context
Across **52,988 attempts**, same-window rankings reached Spearman **0.400** against a required 0.90; byte-identical next-day replays reached **0.78** against 0.99. Label mapping, gaps below the noise floor, and changing rankings for identical inputs explained the failures. If an LLM judge gates agent releases or training data, test it as a measurement instrument first. Repeat identical requests, estimate the noise floor, retain execution records, and run a small pilot; the authors say roughly **2% of call volume** would have revealed both unreachable gates. Changing metrics, sampling, waiting, or switching among four tested providers did not repair reliability on the tested grids. The finding concerns externally observed behavior on shared endpoints; self-hosting helped only while the server was quiet, so it is not a general guarantee.
Context Map
benchmarkresearch#agent-evals#benchmark-integrity#agent-reliabilityGeneric AgentPrepare Coding Session
Uncertainty
Automatically selected from source material; feed7 has not independently tested the claim.
Rate This Item
Personal Note
No note yet. Notes are included in exported bundles.
Related — Every Edge Explained
No approved edges yet for this post.