Llm Evaluation

作者 wshobson46891e7e60da無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

僅含說明AI & Agents
AI 產生的概覽

指導如何用自動指標、人工評審、LLM 作為評審和基準測試來評估 LLM 應用。

功能
說明 LLM 應用的評估策略,涵蓋文字生成、分類與檢索的自動指標、人工評估面向,以及 LLM 作為評審的方法。它提供一個以指標為基礎的評估套件 Python 快速入門範例,對模型執行測試案例並彙總分數。更詳細的模式文件放在配套參考檔案中。
適用情境
適用於系統性地衡量 LLM 應用表現、比較不同模型或提示詞、在部署前發現效能退步、驗證提示詞變更效果,以及建立基準與排查異常模型行為。
執行需求
不附帶腳本,僅為說明性內容。它引用配套檔案 references/details.md,範例程式碼假定使用 Python 與 numpy,以及 BLEU、ROUGE、BERTScore 等指標實作。

LLM Evaluation

Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

When to Use This Skill

  • Measuring LLM application performance systematically
  • Comparing different models or prompts
  • Detecting performance regressions before deployment
  • Validating improvements from prompt changes
  • Building confidence in production systems
  • Establishing baselines and tracking progress over time
  • Debugging unexpected model behavior

Core Evaluation Types

1. Automated Metrics

Fast, repeatable, scalable evaluation using computed scores.

Text Generation:

  • BLEU: N-gram overlap (translation)
  • ROUGE: Recall-oriented (summarization)
  • METEOR: Semantic similarity
  • BERTScore: Embedding-based similarity
  • Perplexity: Language model confidence

Classification:

  • Accuracy: Percentage correct
  • Precision/Recall/F1: Class-specific performance
  • Confusion Matrix: Error patterns
  • AUC-ROC: Ranking quality

Retrieval (RAG):

  • MRR: Mean Reciprocal Rank
  • NDCG: Normalized Discounted Cumulative Gain
  • Precision@K: Relevant in top K
  • Recall@K: Coverage in top K

2. Human Evaluation

Manual assessment for quality aspects difficult to automate.

Dimensions:

  • Accuracy: Factual correctness
  • Coherence: Logical flow
  • Relevance: Answers the question
  • Fluency: Natural language quality
  • Safety: No harmful content
  • Helpfulness: Useful to the user

3. LLM-as-Judge

Use stronger LLMs to evaluate weaker model outputs.

Approaches:

  • Pointwise: Score individual responses
  • Pairwise: Compare two responses
  • Reference-based: Compare to gold standard
  • Reference-free: Judge without ground truth

Quick Start

python
from dataclasses import dataclassfrom typing import Callableimport numpy as np
@dataclassclass Metric:    name: str    fn: Callable
    @staticmethod    def accuracy():        return Metric("accuracy", calculate_accuracy)
    @staticmethod    def bleu():        return Metric("bleu", calculate_bleu)
    @staticmethod    def bertscore():        return Metric("bertscore", calculate_bertscore)
    @staticmethod    def custom(name: str, fn: Callable):        return Metric(name, fn)
class EvaluationSuite:    def __init__(self, metrics: list[Metric]):        self.metrics = metrics
    async def evaluate(self, model, test_cases: list[dict]) -> dict:        results = {m.name: [] for m in self.metrics}
        for test in test_cases:            prediction = await model.predict(test["input"])
            for metric in self.metrics:                score = metric.fn(                    prediction=prediction,                    reference=test.get("expected"),                    context=test.get("context")                )                results[metric.name].append(score)
        return {            "metrics": {k: np.mean(v) for k, v in results.items()},            "raw_scores": results        }
# Usagesuite = EvaluationSuite([    Metric.accuracy(),    Metric.bleu(),    Metric.bertscore(),    Metric.custom("groundedness", check_groundedness)])
test_cases = [    {        "input": "What is the capital of France?",        "expected": "Paris",        "context": "France is a country in Europe. Paris is its capital."    },]
results = await suite.evaluate(model=your_model, test_cases=test_cases)

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

來源與署名

來源:wshobson/agents位於plugins/llm-application-dev/skills/llm-evaluation提交46891e7

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架