Symbolic Equation

作者 lingzhi2279e6c085d65e3無授權條款386 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫7 個月前更新

Discover scientific equations from data using LLM-guided evolutionary search (LLM-SR). Multi-island algorithm with softmax-based cluster sampling, island reset, and LLM-proposed equation mutations. Use for symbolic regression and equation discovery.

AI 產生的概覽

運用 LLM 引導的演化搜尋與多島取樣,從資料中找出可解釋的科學方程式。

功能
引導代理執行 LLM-SR 符號迴歸:先定義包含輸入變數、輸出變數、適應度函式與物理背景的問題規格,再執行多島演化迴圈,由 LLM 提出改良的方程式變體,並加以評估並依表現分群。它透過溫度排程的 softmax 分群取樣與週期性島嶼重置維持多樣性,最後擷取、排序並化簡最佳方程式,並提供物理詮釋。
適用情境
適用於需要從數值資料還原可解釋數學方程式或公式的情境,尤其是能借助領域知識引導搜尋的科學或物理問題。它針對符號迴歸任務,而非配適不透明的黑箱模型。
執行需求
需要一個 LLM 來提出方程式變異,以及具備 numpy 與 scipy(BFGS 或 Adam 參數最佳化)的 Python 環境來評估候選方程式。此技能未附帶指令碼,僅引用一份模式文件,並預期代理自行撰寫並執行方程式程式碼。

Symbolic Equation Discovery

Discover interpretable scientific equations from data using LLM-guided evolutionary search.

Input

  • $0 — Dataset description, variable names, and physical context

References

  • LLM-SR patterns (prompts, evolution, sampling): ~/.claude/skills/symbolic-equation/references/llmsr-patterns.md

Workflow (from LLM-SR)

Step 1: Define Problem Specification

Create a specification with:

  1. Input variables: Physical quantities with types (e.g., x: np.ndarray, v: np.ndarray)
  2. Output variable: Target quantity to predict
  3. Evaluation function: Fitness metric (typically negative MSE with parameter optimization)
  4. Physical context: Domain knowledge to guide equation discovery
python
# Example specification@equation.evolvedef equation(x: np.ndarray, v: np.ndarray, params: np.ndarray) -> np.ndarray:    """Describe the acceleration of a damped nonlinear oscillator."""    return params[0] * x

Step 2: Initialize Multi-Island Buffer

  • Create N islands (default: 10) for population diversity
  • Each island maintains independent clusters of equations
  • Clusters group equations by performance signature

Step 3: Evolutionary Search Loop

Repeat until convergence or max samples:

  1. Select island: Random island selection
  2. Build prompt: Sample top equations from clusters (softmax-weighted by score)
  3. LLM proposes: Generate new equation as improved version
  4. Evaluate: Execute on test data, compute fitness score
  5. Register: Add to island's cluster if valid

Step 4: Prompt Construction

Present previous equations as versioned sequence:

python
def equation_v0(x, v, params):    """Initial version."""    return params[0] * x
def equation_v1(x, v, params):    """Improved version of equation_v0."""    return params[0] * x + params[1] * v
def equation_v2(x, v, params):    """Improved version of equation_v1."""    # LLM completes this

Step 5: Island Reset (Diversity Maintenance)

Periodically (default: every 4 hours):

  1. Sort islands by best score
  2. Reset bottom 50% of islands
  3. Seed each reset island with best equation from a surviving island
  4. Restart cluster sampling temperature

Step 6: Extract Best Equations

After search completes:

  1. Collect best equation from each island
  2. Rank by fitness score
  3. Simplify if possible (algebraic simplification)
  4. Report with physical interpretation

Cluster Sampling

Temperature-scheduled softmax over cluster scores:

temperature = T_init * (1 - (num_programs % period) / period)probabilities = softmax(cluster_scores / temperature)
  • Higher temperature → more exploration
  • Lower temperature → more exploitation of best clusters
  • Within clusters: shorter programs are preferred (Occam's razor)

Rules

  • Equations must use only standard mathematical operations
  • Parameter optimization via scipy BFGS or Adam
  • Fitness = negative MSE (higher is better)
  • Timeout protection for equation evaluation
  • No recursive equations allowed
  • Physical interpretability is preferred over pure fit

Related Skills

  • Upstream: data-analysis, math-reasoning
  • Downstream: paper-writing-section
  • See also: algorithm-design

來源與署名

來源:lingzhi227/agent-research-skills位於skills/symbolic-equation提交9e6c085

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架