Setup

alirezarezvani/claude-skills/engineering/autoresearch-agent/skills/setup

by alirezarezvani19392f7a0826No license27K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 5 weeks ago

Set up a new autoresearch experiment interactively. Collects domain, target file, eval command, metric, direction, and evaluator. Use when the user runs /ar:setup or asks to start optimizing a file with the autoresearch loop.

Instructions onlyAI & Agents
AI-generated overview

Sets up a new autoresearch experiment by collecting domain, target file, eval command, metric, direction, and evaluator.

What it does
This skill configures a new autoresearch experiment, either interactively or from command-line arguments. It gathers parameters such as domain, experiment name, target file, evaluation command, metric, optimization direction, evaluator, and storage scope, then reports the experiment path, branch name, and baseline metric. It can also list existing experiments and available built-in evaluators.
When to use it
Use it when starting a new autoresearch optimization experiment, or when the user runs /ar:setup. It is also used to review existing experiments or see which built-in evaluators are available.
Requirements
Requires the setup_experiment.py script referenced by the skill, a Python runtime, and a working evaluation command for the target file. The skill itself ships no scripts and is instructions only.

/ar:setup — Create New Experiment

Set up a new autoresearch experiment with all required configuration.

Usage

/ar:setup                                    # Interactive mode/ar:setup engineering api-speed src/api.py "pytest bench.py" p50_ms lower/ar:setup --list                             # Show existing experiments/ar:setup --list-evaluators                  # Show available evaluators

What It Does

If arguments provided

Pass them directly to the setup script:

bash
python {skill_path}/scripts/setup_experiment.py \  --domain {domain} --name {name} \  --target {target} --eval "{eval_cmd}" \  --metric {metric} --direction {direction} \  [--evaluator {evaluator}] [--scope {scope}]

If no arguments (interactive mode)

Collect each parameter one at a time:

  1. Domain — Ask: "What domain? (engineering, marketing, content, prompts, custom)"
  2. Name — Ask: "Experiment name? (e.g., api-speed, blog-titles)"
  3. Target file — Ask: "Which file to optimize?" Verify it exists.
  4. Eval command — Ask: "How to measure it? (e.g., pytest bench.py, python evaluate.py)"
  5. Metric — Ask: "What metric does the eval output? (e.g., p50_ms, ctr_score)"
  6. Direction — Ask: "Is lower or higher better?"
  7. Evaluator (optional) — Show built-in evaluators. Ask: "Use a built-in evaluator, or your own?"
  8. Scope — Ask: "Store in project (.autoresearch/) or user (~/.autoresearch/)?"

Then run setup_experiment.py with the collected parameters.

Listing

bash
# Show existing experimentspython {skill_path}/scripts/setup_experiment.py --list
# Show available evaluatorspython {skill_path}/scripts/setup_experiment.py --list-evaluators

Built-in Evaluators

NameMetricUse Case
benchmark_speedp50_ms (lower)Function/API execution time
benchmark_sizesize_bytes (lower)File, bundle, Docker image size
test_pass_ratepass_rate (higher)Test suite pass percentage
build_speedbuild_seconds (lower)Build/compile/Docker build time
memory_usagepeak_mb (lower)Peak memory during execution
llm_judge_contentctr_score (higher)Headlines, titles, descriptions
llm_judge_promptquality_score (higher)System prompts, agent instructions
llm_judge_copyengagement_score (higher)Social posts, ad copy, emails

After Setup

Report to the user:

  • Experiment path and branch name
  • Whether the eval command worked and the baseline metric
  • Suggest: "Run /ar:run {domain}/{name} to start iterating, or /ar:loop {domain}/{name} for autonomous mode."

Source and attribution

Source:alirezarezvani/claude-skillsinengineering/autoresearch-agent/skills/setupat commit19392f7

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal