Sign InOpen Brain
arXivPaperNeeds Review

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

A controlled study encoded authenticated internal-state evidence into unchanged answers, suggesting generated text could carry provenance signals, but not that current models reveal them naturally.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

Researchers forced arithmetic models through one of two discrete internal paths, authenticated the chosen state, and encoded a subtle detectable pattern in otherwise equivalent output. Feed-forward and transformer systems passed **all 128 matched pairs** in public and sealed evaluations.

Practical Implication

For builders, this sketches a future provenance mechanism in which output carries evidence about a causally relevant computation, not just the final answer. Replication across **five feed-forward models** and **three transformers** supports the controlled mechanism.

Agent-Ready Context
Researchers forced arithmetic models through one of two discrete internal paths, authenticated the chosen state, and encoded a subtle detectable pattern in otherwise equivalent output. Feed-forward and transformer systems passed **all 128 matched pairs** in public and sealed evaluations.

For builders, this sketches a future provenance mechanism in which output carries evidence about a causally relevant computation, not just the final answer. Replication across **five feed-forward models** and **three transformers** supports the controlled mechanism.

This is a bounded proof of concept using deliberately trained architectures and an engineered signal. In a separate answer-only transformer experiment, linear probes did not recover a naturally learned intermediate state, so the work does not validate provenance for ordinary model outputs.
Context Map
benchmarksecurity#agent-evals#benchmark-integrity
Uncertainty
This is a bounded proof of concept using deliberately trained architectures and an engineered signal. In a separate answer-only transformer experiment, linear probes did not recover a naturally learned intermediate state, so the work does not validate provenance for ordinary model outputs.