EvalForge Lite

io.github.thejaredchapmanv1.0.1更新於 Oct 1, 2026

Compare text LLMs across OpenRouter, Bedrock, Vertex AI and Foundry with automated grading.

已驗證STDIO僅桌面Data & AnalyticsAI & ML

概覽

AI 產生的概覽

用你自己的提示詞比較 OpenRouter、Bedrock、Vertex AI 和 Foundry 上最多四個文字 LLM,並自動評分。

功能
EvalForge Lite 會把相同的測試提示詞送到最多四個模型、涵蓋四個後端,並用評審模型加上 contains、regex、json_valid、max_length 等規則檢查來為答案評分。它會產生排行榜,包含 0-100 分數、字母等級、延遲、每秒 token 數與預估成本,還能在上傳公司政策後攔截違規的提示詞。九個 MCP 工具涵蓋列出與建議模型、查詢可用性、設定政策、評估單一提示詞、執行比較、列出執行紀錄,以及取得文字或 CSV 報告。
適用情境
當你想根據自己任務的證據、而不是通用基準來挑選模型時,或需要在一次執行中比較同一個模型在兩個平台上的表現時,就很適合。它也適合在決定供應商之前快速比較品質、速度與成本。
執行需求
透過 stdio 在本機執行,通常用 uvx 從 PyPI 套件 evalforge-lite 啟動;網頁版需要 Python 3.10 或更新版本。每個後端都需要對應的憑證:OpenRouter API 金鑰,或區域加上 Bedrock API 金鑰或 AWS 存取金鑰,或 Vertex 專案 ID 與區域加上存取權杖或服務帳戶 JSON,或 Foundry 資源名稱與區域加上 API 金鑰或 Entra ID 權杖。選用環境變數 OPENROUTER_API_KEY 可讓伺服器保存該金鑰,使工具呼叫不必再傳憑證。
安裝前請注意
它會處理真實的供應商憑證,包括 AWS 存取金鑰、服務帳戶 JSON 與 Entra ID 權杖,並把你的提示詞與評分標準送到所選模型供應商與評審後端。執行會在供應商端產生費用;Bedrock、Vertex 與 Foundry 的數字是依型錄價格估算,並非你的雲端帳單。狀態只存在記憶體中,重新啟動即清空,而且應用程式本身不會讀取 .env。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 EvalForge Lite,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

EvalForge Lite

Compare text LLMs side by side. Write a few test prompts, pick up to four models across OpenRouter, Amazon Bedrock, Google Vertex AI and Microsoft Foundry, and EvalForge Lite sends the same prompts to all of them, scores the answers automatically, and shows a leaderboard with letter grades, response time, speed and estimated cost. Use it as a web app or as an MCP server for Claude and other assistants. You bring your own credentials, or the person hosting it keeps them on the server for you.

Documentation site: https://thejaredchapman.github.io/evalforge-lite/

Why use it

  • Test on your own prompts. Choose models from evidence about your tasks, not a generic benchmark.
  • Up to 4 models per run, across 4 backends. X (OpenRouter) and X@bedrock are separate targets, so you can check one model on two platforms in a single run.
  • Automatic grading. A judge model scores each answer against your rubric, and rule checks (contains, regex, json_valid, max_length, available through the API and MCP) add a pass or fail. You get a 0-100 score and a letter grade.
  • A second opinion on every response. Each answer is also evaluated on six criteria (answered, quality, instruction following, completeness, helpfulness, safety) with strengths and weaknesses written out.
  • Speed and cost beside quality. Latency, tokens per second and estimated cost for every model, and a "What matters most?" selector that moves the "Best for ..." badge without a new run.
  • Policy gate. Upload a company policy and prompts that violate it are blocked before any model is called. If the check itself fails, the prompt is blocked.
  • Reports. Download a PDF or a CSV for any of your last five runs.
  • No accounts, no database. Credentials are used for one request and not stored. Nothing is written to disk.

Quick start

1. Run the web app on your computer

Requires Python 3.10 or newer (3.12 recommended; download from https://www.python.org/downloads/) and git.

git clone https://github.com/thejaredchapman/evalforge-lite.gitcd evalforge-litepython3.12 -m venv venvsource venv/bin/activatepip install -r requirements.txtpython app.py

Open http://localhost:8000, paste a key for at least one backend (an OpenRouter API key is the quickest: https://openrouter.ai/workspaces/default/keys), add a test case, pick two to four models, and click Run comparison. Runs are limited to 3 per 8 hours per browser session. Full walkthrough: Getting started.

2. Use it from Claude (MCP server)

With uv installed:

uvx evalforge-lite

Add it to Claude Code in one line:

claude mcp add evalforge-lite -- uvx evalforge-lite

Or install the Claude Code plugin, which bundles the same server:

claude plugin marketplace add thejaredchapman/evalforge-liteclaude plugin install evalforge-lite@evalforge

Then ask your assistant to compare models. It gets 9 tools: list_models, suggest_models, list_availability, set_policy, evaluate_prompt, run_comparison, list_runs, get_report, get_report_csv. Details, Claude Desktop config and credential shapes: MCP server.

3. Host it for other people

Deploy with the included render.yaml (gunicorn, one worker) or any host that can run gunicorn --workers 1 --threads 4 --bind 0.0.0.0:$PORT app:app. By default every visitor supplies their own key. Optionally keep provider keys on the server with environment variables and a shared daily cap (50 per 24 hours by default). Keep it at one worker: all state is in memory per process. Full guide: Hosting and server-side keys.

Good to know

  • The app does not read .env by itself. To use values from it, run set -a; source .env; set +a before python app.py.
  • Each run is limited to 4 models, and each browser session gets 3 runs per 8 hours.
  • All state lives in memory and is cleared when the server restarts. See Privacy and limits.
  • Upgrading from an older version and reading the CSV or API fields? See the notes in Troubleshooting and FAQ.

Backends and credentials

BackendWhat you provide
OpenRouterOne API key
Amazon BedrockA region, plus a Bedrock API key or AWS access keys (optional session token)
Google Vertex AIA project id and region, plus an access token or service-account JSON
Microsoft FoundryA resource name and region, plus an API key or Entra ID access token

A separate judge backend setting chooses where the judge and policy gate run. Bedrock, Vertex and Foundry costs are estimates from catalog prices, not your cloud bill. See Backends and credentials.

Documentation

PageWhat is in it
OverviewWhat it is, who it is for, the three ways to use it
Getting startedInstall, run, and your first comparison
Web app guideEvery part of the screen, in order
Comparing modelsReading metrics, grades, evaluation, cost and their limits
Backends and credentialsKeys, regions, X@backend targets
MCP serverInstall paths, all 9 tools, example prompts
Hosting and server-side keysDeploying for others, operator-held keys, daily cap
Troubleshooting and FAQCommon messages, fixes, and notes on CSV/API field changes
Privacy and limitsWhat data goes where, what is stored, every limit

Test

pytest tests/ -v

Every model and HTTP call is mocked, so the suite needs no API key and makes no network calls.

Contributing

Contributions are welcome: bug reports, model-catalog updates, new checks, docs, and new backends. See CONTRIBUTING.md for setup, tests and the pull request process. When the app shows an error, the popup's Report an issue on GitHub button opens a pre-filled bug report.

License

MIT

來源:README.md,提交 b774029

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v1.0.1最新Oct 1, 2026