Llm Evaluation

作者 wshobson46891e7e60da无许可证收录于 2026年10月8日更新于 2026年10月8日

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

仅含说明AI & Agents
AI 生成的概览

指导如何用自动指标、人工评审、LLM 作为评审和基准测试来评估 LLM 应用。

功能
说明 LLM 应用的评估策略,涵盖文本生成、分类和检索的自动指标、人工评估维度,以及 LLM 作为评审的方法。它给出一个基于指标的评估套件的 Python 快速入门示例,对模型运行测试用例并汇总分数。更详细的模式文档放在配套参考文件中。
适用场景
适用于系统性地衡量 LLM 应用表现、比较不同模型或提示词、在部署前发现性能回退、验证提示词改动效果,以及建立基线和排查异常模型行为。
运行要求
不附带脚本,仅为说明性内容。它引用配套文件 references/details.md,示例代码假定使用 Python 与 numpy,以及 BLEU、ROUGE、BERTScore 等指标实现。

LLM Evaluation

Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

When to Use This Skill

  • Measuring LLM application performance systematically
  • Comparing different models or prompts
  • Detecting performance regressions before deployment
  • Validating improvements from prompt changes
  • Building confidence in production systems
  • Establishing baselines and tracking progress over time
  • Debugging unexpected model behavior

Core Evaluation Types

1. Automated Metrics

Fast, repeatable, scalable evaluation using computed scores.

Text Generation:

  • BLEU: N-gram overlap (translation)
  • ROUGE: Recall-oriented (summarization)
  • METEOR: Semantic similarity
  • BERTScore: Embedding-based similarity
  • Perplexity: Language model confidence

Classification:

  • Accuracy: Percentage correct
  • Precision/Recall/F1: Class-specific performance
  • Confusion Matrix: Error patterns
  • AUC-ROC: Ranking quality

Retrieval (RAG):

  • MRR: Mean Reciprocal Rank
  • NDCG: Normalized Discounted Cumulative Gain
  • Precision@K: Relevant in top K
  • Recall@K: Coverage in top K

2. Human Evaluation

Manual assessment for quality aspects difficult to automate.

Dimensions:

  • Accuracy: Factual correctness
  • Coherence: Logical flow
  • Relevance: Answers the question
  • Fluency: Natural language quality
  • Safety: No harmful content
  • Helpfulness: Useful to the user

3. LLM-as-Judge

Use stronger LLMs to evaluate weaker model outputs.

Approaches:

  • Pointwise: Score individual responses
  • Pairwise: Compare two responses
  • Reference-based: Compare to gold standard
  • Reference-free: Judge without ground truth

Quick Start

python
from dataclasses import dataclassfrom typing import Callableimport numpy as np
@dataclassclass Metric:    name: str    fn: Callable
    @staticmethod    def accuracy():        return Metric("accuracy", calculate_accuracy)
    @staticmethod    def bleu():        return Metric("bleu", calculate_bleu)
    @staticmethod    def bertscore():        return Metric("bertscore", calculate_bertscore)
    @staticmethod    def custom(name: str, fn: Callable):        return Metric(name, fn)
class EvaluationSuite:    def __init__(self, metrics: list[Metric]):        self.metrics = metrics
    async def evaluate(self, model, test_cases: list[dict]) -> dict:        results = {m.name: [] for m in self.metrics}
        for test in test_cases:            prediction = await model.predict(test["input"])
            for metric in self.metrics:                score = metric.fn(                    prediction=prediction,                    reference=test.get("expected"),                    context=test.get("context")                )                results[metric.name].append(score)
        return {            "metrics": {k: np.mean(v) for k, v in results.items()},            "raw_scores": results        }
# Usagesuite = EvaluationSuite([    Metric.accuracy(),    Metric.bleu(),    Metric.bertscore(),    Metric.custom("groundedness", check_groundedness)])
test_cases = [    {        "input": "What is the capital of France?",        "expected": "Paris",        "context": "France is a country in Europe. Paris is its capital."    },]
results = await suite.evaluate(model=your_model, test_cases=test_cases)

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

来源与署名

来源:wshobson/agents位于plugins/llm-application-dev/skills/llm-evaluation提交46891e7

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架