arXivPaperNeeds Review
Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
The same model identifier produced sharply different judgments across API and web deployments. Treat model, interface, system configuration, and date as one versioned dependency.
arXiv · Jul 24, 2026
Source Summary
Researchers tested **four model families** from **October 2025 to February 2026** on a contested pseudo-scientific claim. Grok Fast scored it 70–75 versus 15–40 for the other families, while control prompts did not show the same gap.
Practical Implication
For research agents, do not treat a model ID as a stable epistemic contract. Pin the deployment channel, record dates and configuration, and rerun domain-specific validation after silent service changes.
Agent-Ready Context
Researchers tested **four model families** from **October 2025 to February 2026** on a contested pseudo-scientific claim. Grok Fast scored it 70–75 versus 15–40 for the other families, while control prompts did not show the same gap. For research agents, do not treat a model ID as a stable epistemic contract. Pin the deployment channel, record dates and configuration, and rerun domain-specific validation after silent service changes. The starkest mismatch was the same Grok identifier scoring **75 through the API and 5.5 on the web** three months later. This is a narrow case study, so it demonstrates deployment sensitivity rather than broad comparative model quality.
Context Map
benchmarkresearch#model-selection#agent-reliability#benchmark-integrityUncertainty
The starkest mismatch was the same Grok identifier scoring **75 through the API and 5.5 on the web** three months later. This is a narrow case study, so it demonstrates deployment sensitivity rather than broad comparative model quality.