Autoresearchclaw Autonomous Research

作者 reason-machines2384a003145a无许可证83 个星标收录于 2026年10月8日更新于 2026年10月8日仓库3个月前更新

Fully autonomous research pipeline that turns a topic idea into a complete academic paper with real citations, experiments, and conference-ready LaTeX.

AI 生成的概览

运行自主的 23 阶段流水线,把研究主题变成带真实引用和 LaTeX 的完整学术论文。

功能
AutoResearchClaw 引导智能体走完 23 个研究阶段:主题界定、从 arXiv 和 Semantic Scholar 收集文献、生成假设、设计并在沙箱中执行实验、分析结果、撰写论文、同行评审以及引用核验。产出包括 Markdown 论文草稿、可直接投稿的 LaTeX、BibTeX 文件、核验报告、评审意见、实验运行记录和图表。运行可以续跑、只执行指定阶段,或以编程方式批量执行。
适用场景
当你希望从一句自然语言研究想法直接得到完整论文草稿,而不想手动检索文献或编排实验时使用。它适合自动化或批量论文生成、用配置文件进行的可复现研究运行,以及续跑或查看此前的运行记录。
运行要求
需要 Python 3.11+ 以及从其代码仓库安装的 AutoResearchClaw 包;需要 LLM 服务凭据(OpenAI 或 OpenRouter API 密钥),或 claude、codex、gemini 等 ACP 智能体命令行工具;需要网络访问以查询 arXiv 和 Semantic Scholar;需要 LaTeX 发行版来编译输出;OpenClaw 桥接功能为可选。该技能本身不附带脚本。

AutoResearchClaw — Autonomous Research Pipeline

Skill by ara.so — Daily 2026 Skills collection.

AutoResearchClaw is a fully autonomous 23-stage research pipeline that takes a natural language topic and produces a complete academic paper: real arXiv/Semantic Scholar citations, sandboxed experiments, statistical analysis, multi-agent peer review, and conference-ready LaTeX (NeurIPS/ICML/ICLR). No hallucinated references. No human babysitting.


Installation

bash
# Clone and installgit clone https://github.com/aiming-lab/AutoResearchClaw.gitcd AutoResearchClawpython3 -m venv .venv && source .venv/bin/activatepip install -e .
# Verify CLI is availableresearchclaw --help

Requirements: Python 3.11+


Configuration

bash
cp config.researchclaw.example.yaml config.arc.yaml

Minimum config (config.arc.yaml)

yaml
project:  name: "my-research"
research:  topic: "Your research topic here"
llm:  provider: "openai"  base_url: "https://api.openai.com/v1"  api_key_env: "OPENAI_API_KEY"  primary_model: "gpt-4o"  fallback_models: ["gpt-4o-mini"]
experiment:  mode: "sandbox"  sandbox:    python_path: ".venv/bin/python"
bash
export OPENAI_API_KEY="$YOUR_OPENAI_KEY"

OpenRouter config (200+ models)

yaml
llm:  provider: "openrouter"  api_key_env: "OPENROUTER_API_KEY"  primary_model: "anthropic/claude-3.5-sonnet"  fallback_models:    - "google/gemini-pro-1.5"    - "meta-llama/llama-3.1-70b-instruct"
bash
export OPENROUTER_API_KEY="$YOUR_OPENROUTER_KEY"

ACP (Agent Client Protocol) — no API key needed

yaml
llm:  provider: "acp"  acp:    agent: "claude"   # or: codex, gemini, opencode, kimi    cwd: "."

The agent CLI (e.g. claude) handles its own authentication.

OpenClaw bridge (optional advanced capabilities)

yaml
openclaw_bridge:  use_cron: true              # Scheduled research runs  use_message: true           # Progress notifications  use_memory: true            # Cross-session knowledge persistence  use_sessions_spawn: true    # Parallel sub-sessions  use_web_fetch: true         # Live web search in literature review  use_browser: false          # Browser-based paper collection

Key CLI Commands

bash
# Basic run — fully autonomous, no promptsresearchclaw run --topic "Your research idea" --auto-approve
# Run with explicit config fileresearchclaw run --config config.arc.yaml --topic "Mixture-of-experts routing efficiency" --auto-approve
# Run with topic defined in config (omit --topic flag)researchclaw run --config config.arc.yaml --auto-approve
# Interactive mode — pauses at gate stages for approvalresearchclaw run --config config.arc.yaml --topic "Your topic"
# Check pipeline status / resume a runresearchclaw status --run-id rc-20260315-120000-abc123
# List past runsresearchclaw list

Gate stages (5, 9, 20) pause for human approval in interactive mode. Pass --auto-approve to skip all gates.


Python API

python
from researchclaw.pipeline import Runnerfrom researchclaw.config import load_config
# Load config and runconfig = load_config("config.arc.yaml")config.research.topic = "Efficient attention mechanisms for long-context LLMs"config.auto_approve = True
runner = Runner(config)result = runner.run()
# Access outputsprint(result.artifact_dir)          # artifacts/rc-YYYYMMDD-HHMMSS-<hash>/print(result.deliverables_dir)      # .../deliverables/print(result.paper_draft_path)      # .../deliverables/paper_draft.mdprint(result.latex_path)            # .../deliverables/paper.texprint(result.bibtex_path)           # .../deliverables/references.bibprint(result.verification_report)  # .../deliverables/verification_report.json
python
# Run specific stages onlyfrom researchclaw.pipeline import Runner, StageRange
runner = Runner(config)result = runner.run(stages=StageRange(start="LITERATURE_COLLECT", end="KNOWLEDGE_EXTRACT"))
python
# Access knowledge base after a runfrom researchclaw.knowledge import KnowledgeBase
kb = KnowledgeBase.load(result.artifact_dir)findings = kb.get("findings")literature = kb.get("literature")decisions = kb.get("decisions")

Output Structure

After a run, all outputs land in artifacts/rc-YYYYMMDD-HHMMSS-<hash>/:

artifacts/rc-20260315-120000-abc123/├── deliverables/│   ├── paper_draft.md          # Full academic paper (Markdown)│   ├── paper.tex               # Conference-ready LaTeX│   ├── references.bib          # Real BibTeX — auto-pruned to inline citations│   ├── verification_report.json # 4-layer citation integrity report│   └── reviews.md              # Multi-agent peer review├── experiment_runs/│   ├── run_001/│   │   ├── code/               # Generated experiment code│   │   ├── results.json        # Structured metrics│   │   └── sandbox_output.txt  # Execution logs├── charts/│   └── *.png                   # Auto-generated comparison charts├── evolution/│   └── lessons.json            # Self-learning lessons for future runs└── knowledge_base/    ├── decisions.json    ├── experiments.json    ├── findings.json    ├── literature.json    ├── questions.json    └── reviews.json

Pipeline Stages Reference

PhaseStage #NameNotes
A1TOPIC_INITParse and scope research topic
A2PROBLEM_DECOMPOSEBreak into sub-problems
B3SEARCH_STRATEGYBuild search queries
B4LITERATURE_COLLECTReal API calls to arXiv + Semantic Scholar
B5LITERATURE_SCREENGate — approve/reject literature
B6KNOWLEDGE_EXTRACTExtract structured knowledge
C7SYNTHESISSynthesize findings
C8HYPOTHESIS_GENMulti-agent debate to form hypotheses
D9EXPERIMENT_DESIGNGate — approve/reject design
D10CODE_GENERATIONGenerate experiment code
D11RESOURCE_PLANNINGGPU/MPS/CPU auto-detection
E12EXPERIMENT_RUNSandboxed execution
E13ITERATIVE_REFINESelf-healing on failure
F14RESULT_ANALYSISMulti-agent analysis
F15RESEARCH_DECISIONPROCEED / REFINE / PIVOT
G16PAPER_OUTLINEStructure paper
G17PAPER_DRAFTWrite full paper
G18PEER_REVIEWEvidence-consistency check
G19PAPER_REVISIONIncorporate review feedback
H20QUALITY_GATEGate — final approval
H21KNOWLEDGE_ARCHIVESave lessons to KB
H22EXPORT_PUBLISHEmit LaTeX + BibTeX
H23CITATION_VERIFY4-layer anti-hallucination check

Common Patterns

Pattern: Quick paper on a topic

bash
export OPENAI_API_KEY="$OPENAI_API_KEY"researchclaw run \  --topic "Self-supervised learning for protein structure prediction" \  --auto-approve

Pattern: Reproducible run with full config

yaml
# config.arc.yamlproject:  name: "protein-ssl-research"
research:  topic: "Self-supervised learning for protein structure prediction"
llm:  provider: "openai"  api_key_env: "OPENAI_API_KEY"  primary_model: "gpt-4o"  fallback_models: ["gpt-4o-mini"]
experiment:  mode: "sandbox"  sandbox:    python_path: ".venv/bin/python"  max_iterations: 3  timeout_seconds: 300
bash
researchclaw run --config config.arc.yaml --auto-approve

Pattern: Use Claude via OpenRouter for best reasoning

bash
export OPENROUTER_API_KEY="$OPENROUTER_API_KEY"
cat > config.arc.yaml << 'EOF'project:  name: "my-research"llm:  provider: "openrouter"  api_key_env: "OPENROUTER_API_KEY"  primary_model: "anthropic/claude-3.5-sonnet"  fallback_models: ["google/gemini-pro-1.5"]experiment:  mode: "sandbox"  sandbox:    python_path: ".venv/bin/python"EOF
researchclaw run --config config.arc.yaml \  --topic "Efficient KV cache compression for transformer inference" \  --auto-approve

Pattern: Resume after a failed run

bash
# List runs to find the run IDresearchclaw list
# Resume from last completed stageresearchclaw run --resume rc-20260315-120000-abc123

Pattern: Programmatic batch research

python
import asynciofrom researchclaw.pipeline import Runnerfrom researchclaw.config import load_config
topics = [    "LoRA fine-tuning on limited hardware",    "Speculative decoding for LLM inference",    "Flash attention variants comparison",]
config = load_config("config.arc.yaml")config.auto_approve = True
for topic in topics:    config.research.topic = topic    runner = Runner(config)    result = runner.run()    print(f"[{topic}] → {result.deliverables_dir}")

Pattern: OpenClaw one-liner (if using OpenClaw agent)

Share the repo URL with OpenClaw, then say:"Research mixture-of-experts routing efficiency"

OpenClaw auto-reads RESEARCHCLAW_AGENTS.md, clones, installs, configures, and runs the full pipeline.


Compile the LaTeX Output

bash
# Navigate to deliverablescd artifacts/rc-*/deliverables/
# Compile (requires a LaTeX distribution)pdflatex paper.texbibtex paperpdflatex paper.texpdflatex paper.tex
# Or upload paper.tex + references.bib directly to Overleaf

Troubleshooting

researchclaw: command not found

bash
# Make sure the venv is active and package is installedsource .venv/bin/activatepip install -e .which researchclaw

API key errors

bash
# Verify env var is setecho $OPENAI_API_KEY# Should print your key (not empty)
# Set it explicitly for the sessionexport OPENAI_API_KEY="sk-..."

Experiment sandbox failures

The pipeline self-heals at Stage 13 (ITERATIVE_REFINE). If it keeps failing:

yaml
# Increase timeout and iterations in configexperiment:  max_iterations: 5  timeout_seconds: 600  sandbox:    python_path: ".venv/bin/python"

Citation hallucination warnings

Stage 23 (CITATION_VERIFY) runs a 4-layer check. If references are pruned:

  • This is expected behaviour — fake citations are removed automatically
  • Check verification_report.json for details on which citations were rejected and why

PIVOT loop running indefinitely

Stage 15 (RESEARCH_DECISION) may pivot multiple times. To cap iterations:

yaml
research:  max_pivots: 2  max_refines: 3

LaTeX compilation errors

bash
# Check for missing packagespdflatex paper.tex 2>&1 | grep "File.*not found"
# Install missing packages (TeX Live)tlmgr install <package-name>

Out of memory during experiments

yaml
# Force CPU mode in configexperiment:  sandbox:    device: "cpu"    max_memory_gb: 4

Key Concepts

  • PIVOT/REFINE Loop: Stage 15 autonomously decides PROCEED, REFINE (tweak params), or PIVOT (new hypothesis direction). All artifacts are versioned.
  • Multi-Agent Debate: Stages 8, 14, 18 use structured multi-perspective debate — not a single LLM pass.
  • Self-Learning: Each run extracts lessons with 30-day time decay. Future runs on similar topics benefit from past mistakes.
  • Sentinel Watchdog: Background monitor detects NaN/Inf in results, checks paper-evidence consistency, scores citation relevance, and guards against fabrication throughout the run.
  • 4-Layer Citation Verification: arXiv lookup → CrossRef lookup → DataCite lookup → LLM relevance scoring. A citation must pass all layers to survive.

来源与署名

来源:reason-machines/trending-skills位于skills/autoresearchclaw-autonomous-research提交2384a00

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架