# Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

Source: [arXiv](https://arxiv.org/abs/2608.16868v1)  
Feed7 permalink: https://feed7.dev/p/2608-16868v1-00830as  
Published: 2026-08-17T17:50:04.000Z  
Trust: Needs Review (needs_review)

## Why Included

A controlled study encoded authenticated internal-state evidence into unchanged answers, suggesting generated text could carry provenance signals, but not that current models reveal them naturally.

## Source Summary

Researchers forced arithmetic models through one of two discrete internal paths, authenticated the chosen state, and encoded a subtle detectable pattern in otherwise equivalent output. Feed-forward and transformer systems passed **all 128 matched pairs** in public and sealed evaluations.

## Practical Implication

For builders, this sketches a future provenance mechanism in which output carries evidence about a causally relevant computation, not just the final answer. Replication across **five feed-forward models** and **three transformers** supports the controlled mechanism.

## Agent-Ready Context

Researchers forced arithmetic models through one of two discrete internal paths, authenticated the chosen state, and encoded a subtle detectable pattern in otherwise equivalent output. Feed-forward and transformer systems passed **all 128 matched pairs** in public and sealed evaluations.

For builders, this sketches a future provenance mechanism in which output carries evidence about a causally relevant computation, not just the final answer. Replication across **five feed-forward models** and **three transformers** supports the controlled mechanism.

This is a bounded proof of concept using deliberately trained architectures and an engineered signal. In a separate answer-only transformer experiment, linear probes did not recover a naturally learned intermediate state, so the work does not validate provenance for ordinary model outputs.

## Connected Context

Feed7 judgment across 479 accumulated Signals:

This narrows provenance claims from inspecting outputs or probing naturally learned representations to deliberately training an authenticated causal state into the output. The controlled success shows that computation-path evidence can be carried, while the failed answer-only probe warns that ordinary models cannot yet be assumed to expose such provenance naturally.

- [What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models](https://feed7.dev/p/2608-16852v1-0580uyd) — The compliance audit shows probes can appear accurate without reading the governing rule; this work contrasts that failure with an engineered signal tied to a forced, authenticated internal path.
- [QuoteBench: How Matched Scores Can Hide Command-Path Failures](https://feed7.dev/p/2608-13547v1-130h6xd) — QuoteBench shows final scores can conceal which command path produced them; computational provenance sketches a complementary way to attach evidence about a causally relevant internal path to the output itself.
- [The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping](https://feed7.dev/p/2608-06361v1-1n3dr85) — Both reject answer-only confidence: the video study requires event-level traces, while this work tests whether evidence of an intermediate computational state can travel with the answer.

## Context Map

- Layer: benchmark
- Domains: security
- Topics: agent-evals, benchmark-integrity

## Uncertainty

- This is a bounded proof of concept using deliberately trained architectures and an engineered signal. In a separate answer-only transformer experiment, linear probes did not recover a naturally learned intermediate state, so the work does not validate provenance for ordinary model outputs.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
