
EvalForge Lite
io.github.thejaredchapmanv1.0.1更新于 Oct 1, 2026
Compare text LLMs across OpenRouter, Bedrock, Vertex AI and Foundry with automated grading.
概览
用你自己的提示词对比 OpenRouter、Bedrock、Vertex AI 和 Foundry 上最多四个文本大模型,并自动评分。
- 功能
- EvalForge Lite 把相同的测试提示词发送到最多四个模型、覆盖四个后端,并用评审模型加上 contains、regex、json_valid、max_length 等规则检查来给答案打分。它输出排行榜,包含 0-100 分数、字母等级、延迟、每秒 token 数和预估成本,还能在上传公司政策后拦截违规提示词。九个 MCP 工具涵盖列出与推荐模型、查询可用性、设置政策、评估单条提示词、运行对比、列出运行记录,以及获取文本或 CSV 报告。
- 适用场景
- 当你希望基于自己任务的证据而不是通用基准来选择模型时使用它,或者需要在一次运行中比较同一模型在两个平台上的表现。它也适合在选定供应商之前快速对比质量、速度和成本。
- 运行要求
- 通过 stdio 在本地运行,通常用 uvx 从 PyPI 包 evalforge-lite 启动;网页版需要 Python 3.10 或更高版本。每个后端都需要相应凭据:OpenRouter API 密钥,或区域加 Bedrock API 密钥或 AWS 访问密钥,或 Vertex 项目 ID 与区域加访问令牌或服务账号 JSON,或 Foundry 资源名与区域加 API 密钥或 Entra ID 令牌。可选环境变量 OPENROUTER_API_KEY 可让服务器保存该密钥,使工具调用无需再传凭据。
安装
在 SourceWeft 中
- 打开 控制台中的 EvalForge Lite,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
EvalForge Lite
Compare text LLMs side by side. Write a few test prompts, pick up to four models across OpenRouter, Amazon Bedrock, Google Vertex AI and Microsoft Foundry, and EvalForge Lite sends the same prompts to all of them, scores the answers automatically, and shows a leaderboard with letter grades, response time, speed and estimated cost. Use it as a web app or as an MCP server for Claude and other assistants. You bring your own credentials, or the person hosting it keeps them on the server for you.
Documentation site: https://thejaredchapman.github.io/evalforge-lite/
Why use it
- Test on your own prompts. Choose models from evidence about your tasks, not a generic benchmark.
- Up to 4 models per run, across 4 backends.
X(OpenRouter) andX@bedrockare separate targets, so you can check one model on two platforms in a single run. - Automatic grading. A judge model scores each answer against your rubric, and rule checks (
contains,regex,json_valid,max_length, available through the API and MCP) add a pass or fail. You get a 0-100 score and a letter grade. - A second opinion on every response. Each answer is also evaluated on six criteria (answered, quality, instruction following, completeness, helpfulness, safety) with strengths and weaknesses written out.
- Speed and cost beside quality. Latency, tokens per second and estimated cost for every model, and a "What matters most?" selector that moves the "Best for ..." badge without a new run.
- Policy gate. Upload a company policy and prompts that violate it are blocked before any model is called. If the check itself fails, the prompt is blocked.
- Reports. Download a PDF or a CSV for any of your last five runs.
- No accounts, no database. Credentials are used for one request and not stored. Nothing is written to disk.
Quick start
1. Run the web app on your computer
Requires Python 3.10 or newer (3.12 recommended; download from https://www.python.org/downloads/) and git.
Open http://localhost:8000, paste a key for at least one backend (an OpenRouter API key is the quickest: https://openrouter.ai/workspaces/default/keys), add a test case, pick two to four models, and click Run comparison. Runs are limited to 3 per 8 hours per browser session. Full walkthrough: Getting started.
2. Use it from Claude (MCP server)
With uv installed:
Add it to Claude Code in one line:
Or install the Claude Code plugin, which bundles the same server:
Then ask your assistant to compare models. It gets 9 tools: list_models,
suggest_models, list_availability, set_policy, evaluate_prompt,
run_comparison, list_runs, get_report, get_report_csv.
Details, Claude Desktop config and credential shapes: MCP server.
3. Host it for other people
Deploy with the included render.yaml (gunicorn, one worker) or any host that
can run gunicorn --workers 1 --threads 4 --bind 0.0.0.0:$PORT app:app. By
default every visitor supplies their own key. Optionally keep provider keys on
the server with environment variables and a shared daily cap (50 per 24 hours by
default). Keep it at one worker: all state is in memory per process.
Full guide: Hosting and server-side keys.
Good to know
- The app does not read
.envby itself. To use values from it, runset -a; source .env; set +abeforepython app.py. - Each run is limited to 4 models, and each browser session gets 3 runs per 8 hours.
- All state lives in memory and is cleared when the server restarts. See Privacy and limits.
- Upgrading from an older version and reading the CSV or API fields? See the notes in Troubleshooting and FAQ.
Backends and credentials
A separate judge backend setting chooses where the judge and policy gate run. Bedrock, Vertex and Foundry costs are estimates from catalog prices, not your cloud bill. See Backends and credentials.
Documentation
Test
Every model and HTTP call is mocked, so the suite needs no API key and makes no network calls.
Contributing
Contributions are welcome: bug reports, model-catalog updates, new checks, docs, and new backends. See CONTRIBUTING.md for setup, tests and the pull request process. When the app shows an error, the popup's Report an issue on GitHub button opens a pre-filled bug report.
License
来源:README.md,提交 b774029
工具
0版本历史
1- v1.0.1最新Oct 1, 2026


