Sign InOpen Brain
GitHubGitHub RepoNeeds Review

Imbad0202/academic-research-skills

ARS packages research, writing, review, and citation checks as agent skills with explicit human gates; its strongest design lesson is to bound what automated integrity checks can prove.

GitHub · Trending today
Open Source Open MarkdownOpen JSON
Source Summary

Academic Research Skills packages research, writing, review, revision, and finalization for Claude Code, with a separate Codex distribution. **v3.8** adds opt-in claim audits, five blocking warning classes, and calibration thresholds of **FNR below 0.15** and **FPR below 0.10**.

Practical Implication

Reuse its pattern of explicit integrity gates, provenance records, scoped data access, and human checkpoints when building research agents. Keep evidence gathering and mechanical validation delegated while reserving question choice, methods, interpretation, and final claims for people.

Agent-Ready Context
Academic Research Skills packages research, writing, review, revision, and finalization for Claude Code, with a separate Codex distribution. **v3.8** adds opt-in claim audits, five blocking warning classes, and calibration thresholds of **FNR below 0.15** and **FPR below 0.10**.

Reuse its pattern of explicit integrity gates, provenance records, scoped data access, and human checkpoints when building research agents. Keep evidence gathering and mechanical validation delegated while reserving question choice, methods, interpretation, and final claims for people.

Its checks may be sampled or model-mediated. ARS cannot establish that procedures occurred, raw data is authentic, or results reproduce; the project explicitly warns that a consistently reported fabrication can pass.
Connected Context · Feed7 Judgment

This turns the candidates’ general case for specialized research agents and evidence gates into a concrete integrity-oriented workflow with explicit warning classes and calibration targets. It also draws a firmer boundary around assurance: structured review can improve consistency and traceability, but cannot verify that experiments occurred, data are authentic, or findings reproduce, so human ownership and external evidence remain necessary.

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard DocumentsBoth favor taxonomy-driven specialist review over a generic pass; the benchmark supplies domain-specific performance evidence, while ARS adds integrity gates and explicit calibration thresholds.Claude Science, an AI workbench for scientists, is now availableClaude Science demonstrates a similar coordinator-specialist-reviewer structure at workbench scale, while ARS more explicitly reserves research choices and final claims for humans.Don't Build Agents You Can't Answer For — Addy OsmaniARS operationalizes the call for evidence-backed accountability through provenance, validation, warnings, and human checkpoints, while acknowledging that those records cannot prove underlying research authenticity.What Does Done Even Mean? Agents and Paperclip's Liveness Model - Dotta, PaperclipIts blocking warnings and human finalization boundary reinforce the idea that completion requires evidence, verification, authority, and accepted residual risk rather than an agent declaring itself done.
Context Map
agentresearch#skills#multi-agent#agent-reliability
Uncertainty
Its checks may be sampled or model-mediated. ARS cannot establish that procedures occurred, raw data is authentic, or results reproduce; the project explicitly warns that a consistently reported fabrication can pass.