Sign InOpen Brain
arXivPaperNeeds Review

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

The same model identifier produced sharply different judgments across API and web deployments. Treat model, interface, system configuration, and date as one versioned dependency.

arXiv · Jul 24, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Researchers tested **four model families** from **October 2025 to February 2026** on a contested pseudo-scientific claim. Grok Fast scored it 70–75 versus 15–40 for the other families, while control prompts did not show the same gap.

Practical Implication

For research agents, do not treat a model ID as a stable epistemic contract. Pin the deployment channel, record dates and configuration, and rerun domain-specific validation after silent service changes.

Agent-Ready Context
Researchers tested **four model families** from **October 2025 to February 2026** on a contested pseudo-scientific claim. Grok Fast scored it 70–75 versus 15–40 for the other families, while control prompts did not show the same gap.

For research agents, do not treat a model ID as a stable epistemic contract. Pin the deployment channel, record dates and configuration, and rerun domain-specific validation after silent service changes.

The starkest mismatch was the same Grok identifier scoring **75 through the API and 5.5 on the web** three months later. This is a narrow case study, so it demonstrates deployment sensitivity rather than broad comparative model quality.
Context Map
benchmarkresearch#model-selection#agent-reliability#benchmark-integrity
Uncertainty
The starkest mismatch was the same Grok identifier scoring **75 through the API and 5.5 on the web** three months later. This is a narrow case study, so it demonstrates deployment sensitivity rather than broad comparative model quality.