AutoResearchClaw — Autonomous Research Pipeline
Skill by ara.so — Daily 2026 Skills collection.
AutoResearchClaw is a fully autonomous 23-stage research pipeline that takes a natural language topic and produces a complete academic paper: real arXiv/Semantic Scholar citations, sandboxed experiments, statistical analysis, multi-agent peer review, and conference-ready LaTeX (NeurIPS/ICML/ICLR). No hallucinated references. No human babysitting.
Installation
Requirements: Python 3.11+
Configuration
Minimum config (config.arc.yaml)
OpenRouter config (200+ models)
ACP (Agent Client Protocol) — no API key needed
The agent CLI (e.g. claude) handles its own authentication.
OpenClaw bridge (optional advanced capabilities)
Key CLI Commands
Gate stages (5, 9, 20) pause for human approval in interactive mode. Pass --auto-approve to skip all gates.
Python API
Output Structure
After a run, all outputs land in artifacts/rc-YYYYMMDD-HHMMSS-<hash>/:
Pipeline Stages Reference
Common Patterns
Pattern: Quick paper on a topic
Pattern: Reproducible run with full config
Pattern: Use Claude via OpenRouter for best reasoning
Pattern: Resume after a failed run
Pattern: Programmatic batch research
Pattern: OpenClaw one-liner (if using OpenClaw agent)
OpenClaw auto-reads RESEARCHCLAW_AGENTS.md, clones, installs, configures, and runs the full pipeline.
Compile the LaTeX Output
Troubleshooting
researchclaw: command not found
API key errors
Experiment sandbox failures
The pipeline self-heals at Stage 13 (ITERATIVE_REFINE). If it keeps failing:
Citation hallucination warnings
Stage 23 (CITATION_VERIFY) runs a 4-layer check. If references are pruned:
- This is expected behaviour — fake citations are removed automatically
- Check
verification_report.jsonfor details on which citations were rejected and why
PIVOT loop running indefinitely
Stage 15 (RESEARCH_DECISION) may pivot multiple times. To cap iterations:
LaTeX compilation errors
Out of memory during experiments
Key Concepts
- PIVOT/REFINE Loop: Stage 15 autonomously decides PROCEED, REFINE (tweak params), or PIVOT (new hypothesis direction). All artifacts are versioned.
- Multi-Agent Debate: Stages 8, 14, 18 use structured multi-perspective debate — not a single LLM pass.
- Self-Learning: Each run extracts lessons with 30-day time decay. Future runs on similar topics benefit from past mistakes.
- Sentinel Watchdog: Background monitor detects NaN/Inf in results, checks paper-evidence consistency, scores citation relevance, and guards against fabrication throughout the run.
- 4-Layer Citation Verification: arXiv lookup → CrossRef lookup → DataCite lookup → LLM relevance scoring. A citation must pass all layers to survive.

