Sign InOpen Brain
arXivPaperNeeds Review

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

A structured multi-agent reviewer closed part of the gap on rule-heavy documents, suggesting explicit taxonomies, specialized skills, and verification beat a single generic review pass.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

GB/T-Bench turns **488 documents** into **7,306 traceable errors** across 25 types. Among 14 models, the strongest scored **0.3280 CMCS**, versus **0.6640** for human experts.

Practical Implication

For document-review agents, encode the review taxonomy as specialized skills and separate global inspection, targeted diagnosis, rule scanning, and verification. GB/T-Reviewer lifted the best CMCS to **0.5094**.

Agent-Ready Context
GB/T-Bench turns **488 documents** into **7,306 traceable errors** across 25 types. Among 14 models, the strongest scored **0.3280 CMCS**, versus **0.6640** for human experts.

For document-review agents, encode the review taxonomy as specialized skills and separate global inspection, targeted diagnosis, rule scanning, and verification. GB/T-Reviewer lifted the best CMCS to **0.5094**.

The benchmark targets Chinese national-standard documents and uses generated counterexamples, so results may not transfer directly to other regulated corpora. The remaining expert gap is still substantial.
Context Map
agentdata#multi-agent#skills#agent-reliability
Uncertainty
The benchmark targets Chinese national-standard documents and uses generated counterexamples, so results may not transfer directly to other regulated corpora. The remaining expert gap is still substantial.