Sign InOpen Brain
arXivPaperNeeds Review

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

TrajDebug tracks whether errors persist, resolve, or cause terminal failure, offering a sharper way to debug long coding-agent runs than flagging every local mistake.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

**TrajDebug** compresses long histories at multiple granularities, identifies errors from evidence, then tracks resolution status and terminal impact. **TrajErrBench** contains **486 manually annotated failed trajectories** from Tau2Bench and SWE-Bench Pro.

Practical Implication

For coding-agent observability, preserve enough trajectory context to distinguish a harmless local mistake from the earliest unresolved error that caused the final failure. That attribution can make remediation more targeted.

Agent-Ready Context
**TrajDebug** compresses long histories at multiple granularities, identifies errors from evidence, then tracks resolution status and terminal impact. **TrajErrBench** contains **486 manually annotated failed trajectories** from Tau2Bench and SWE-Bench Pro.

For coding-agent observability, preserve enough trajectory context to distinguish a harmless local mistake from the earliest unresolved error that caused the final failure. That attribution can make remediation more targeted.

The abstract claims the best overall performance and actionable downstream feedback but gives no scores or baseline details. Code and data are promised for release, so reproducibility cannot be assessed from this material.
Context Map
benchmarkcoding#agent-evals#agent-reliability#observability
Uncertainty
The abstract claims the best overall performance and actionable downstream feedback but gives no scores or baseline details. Code and data are promised for release, so reproducibility cannot be assessed from this material.