{
  "schema_version": "1.1",
  "id": "s13:https://arxiv.org/abs/2609.04198v1",
  "slug": "2609-04198v1-1mype86",
  "url": "https://feed7.dev/p/2609-04198v1-1mype86",
  "title": "Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints",
  "why_included": "A 52,988-request audit found black-box LLM judges too unstable for preregistered gates. Measure repeatability first, and pilot the judge before fixing evaluation thresholds.",
  "summary": "Across **52,988 attempts**, same-window rankings reached Spearman **0.400** against a required 0.90; byte-identical next-day replays reached **0.78** against 0.99. Label mapping, gaps below the noise floor, and changing rankings for identical inputs explained the failures.",
  "practical_implication": "If an LLM judge gates agent releases or training data, test it as a measurement instrument first. Repeat identical requests, estimate the noise floor, retain execution records, and run a small pilot; the authors say roughly **2% of call volume** would have revealed both unreachable gates.",
  "agent_context": "Across **52,988 attempts**, same-window rankings reached Spearman **0.400** against a required 0.90; byte-identical next-day replays reached **0.78** against 0.99. Label mapping, gaps below the noise floor, and changing rankings for identical inputs explained the failures.\n\nIf an LLM judge gates agent releases or training data, test it as a measurement instrument first. Repeat identical requests, estimate the noise floor, retain execution records, and run a small pilot; the authors say roughly **2% of call volume** would have revealed both unreachable gates.\n\nChanging metrics, sampling, waiting, or switching among four tested providers did not repair reliability on the tested grids. The finding concerns externally observed behavior on shared endpoints; self-hosting helped only while the server was quiet, so it is not a general guarantee.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2609.04198v1",
    "published_at": "2026-09-03T17:59:43.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [
    "research"
  ],
  "topics": [
    "agent-evals",
    "benchmark-integrity",
    "agent-reliability"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "Changing metrics, sampling, waiting, or switching among four tested providers did not repair reliability on the tested grids. The finding concerns externally observed behavior on shared endpoints; self-hosting helped only while the server was quiet, so it is not a general guarantee."
  ],
  "connected_context": {
    "meaning": "This raises repeatability to a prerequisite for using black-box LLM judges as release or data gates. It strengthens prior warnings about noisy or biased evaluators with preregistered failure at substantial scale and a cheap pilot strategy, while narrowing the diagnosis: neither provider switching nor routine metric adjustments repaired the tested shared-endpoint instability, and quiet self-hosting is not a general solution.",
    "corpus_size": 691,
    "generated_at": "2026-09-05T10:08:06.960Z",
    "connections": [
      {
        "title": "Phantom Gains: Auditing Self-Improvement Against a Measured Null",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.20290v1",
        "feed7_url": "https://feed7.dev/p/2608-20290v1-13o5gau",
        "reason": "Both require an empirical null before interpreting evaluation differences; this Signal extends that principle from self-improvement transitions to judge rankings and replay stability."
      },
      {
        "title": "Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.11871v1",
        "feed7_url": "https://feed7.dev/p/2607-11871v1-17vejh0",
        "reason": "The candidate identifies systematic judge bias, whereas this Signal identifies endpoint-level instability; together they show that validity and repeatability are separate requirements."
      },
      {
        "title": "Verifiable Environments for AI in Biology — Kenny Workman, LatchBio",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=3ZMUiFaQ3qg",
        "feed7_url": "https://feed7.dev/p/verifiable-environments-for-ai-in-biology-kenny-workman-latchbio-1vs6y66",
        "reason": "LatchBio’s human validation of brittle graders becomes more consequential here: domain correctness checks cannot rescue a judge whose own outputs fail repeatability tests."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-09-03T17:59:43.000Z",
  "modified_at": "2026-09-03T17:59:43.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2609-04198v1-1mype86",
    "json": "https://feed7.dev/p/2609-04198v1-1mype86.json",
    "markdown": "https://feed7.dev/p/2609-04198v1-1mype86.md"
  }
}