Nonobench

io.github.mauricekleinev1.1.0更新於 Sep 30, 2026

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

已驗證Streamable HTTP可網頁執行AI & MLData & Analytics

概覽

AI 產生的概覽

以唯讀方式存取 Nonobench 基準資料,了解大型語言模型在 5x5 至 20x20 數織謎題上的表現。

功能
Nonobench 是開放原始碼基準,用來評估大型語言模型在數織(Picross)謎題上的推理能力。託管的 MCP 端點提供已發布的基準資料:排行榜、模型、謎題、個別執行紀錄以及解答檢查器,資料與儀表板來自同一批匯出檔案。此端點為無狀態 Streamable HTTP,不需要驗證。
適用情境
當你希望助理查詢數織基準結果、比較各模型在這些謎題上的表現、檢視個別執行紀錄,或依據提示線索檢查某個解答時,可以使用它。它是讀取資料的端點,不能用來執行新的基準測試。
執行需求
需要一個指向託管端點的遠端 MCP 用戶端;未宣告帳號、API 金鑰或標頭。執行基準測試本身是另一回事,需要儲存庫、Bun 執行環境、供視覺化介面使用的 Node.js,以及 OpenRouter API 金鑰。
安裝前請注意
此端點為唯讀且不需驗證,因此不會傳送敏感資訊。請注意匯出的原始結果包含提示詞與模型輸出;在本機執行基準測試會透過 OpenRouter 產生費用,涉及的憑證是 OPENROUTER_API_KEY,金鑰本身的額度上限才是支出的硬性上限。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 Nonobench,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。

其他 MCP 客戶端

把它新增到你客戶端的 mcpServers 設定中。

{
  "mcpServers": {
    "nonobench": {
      "type": "http",
      "url": "https://www.nonobench.com/mcp"
    }
  }
}

README

Nonobench

A benchmark suite for evaluating LLM reasoning capabilities on Nonogram (Picross) puzzle solving across different grid sizes. Results are published at nonobench.com.

Built by Maurice Kleine.

What is a Nonogram?

Nonograms (also known as Picross, Griddlers, or Paint by Numbers) are logic puzzles where you fill in cells on a grid based on numeric clues for each row and column. The clues indicate consecutive groups of filled cells, separated by at least one empty cell. Solving these puzzles requires logical deduction and constraint satisfaction - making them an excellent test of LLM reasoning abilities.

Project Structure

nonobench/├── bench/          # Benchmark runner, results database and exporter└── visualizer/     # Next.js dashboard (nonobench.com), also home of the puzzle set

Prerequisites

Quick Start

1. Clone the Repository

bash
git clone https://github.com/mauricekleine/nonobench.gitcd nonobench

2. Set Up Environment Variables

Create a .env file in the bench/ directory (see bench/.env.example):

bash
OPENROUTER_API_KEY=your_openrouter_api_key_here

This is the only variable needed. The visualizer builds without any.

3. Running Benchmarks

bash
cd benchbun installbun run bench                       # prints the plan and exits, no API callsbun run bench --model <name>        # run one model (repeat --model for more)bun run bench --all-missing         # run every configured model with missing workbun run bench --model <name> --sizes 20x20  # opt in to the extended tier

Runs are incremental and append-only: a model/puzzle pair that already has a successful result is never run again, and the database refuses to overwrite it. Failed runs are retried on the next invocation. Results are stored in bench/results.db (SQLite).

Useful flags:

  • --max-cost <usd> stops launching new puzzles once this session's spend reaches the amount. Requests already in flight still finish, so a session can overshoot by up to --parallel requests per selected model. Use the OpenRouter key's own limit as the hard ceiling.
  • --parallel <n> sets concurrent requests per model (default 10); lower it for rate-limited providers.
  • --limit <n> runs only the first n puzzles of each size, for pilots against a scratch database (NONOBENCH_DB=/tmp/copy.db).

Output modes

New runs ask for the answer as strict structured output (a JSON schema, only routed to endpoints that enforce it), so models cannot wrap the grid in prose. For a few models the schema-enforcing endpoints measurably hurt answers; those run in text mode instead (outputMode: "text" in bench/constants.ts, chosen by a 5x5 A/B with the benchmark prompt). Every run records its mode, and grading is identical for both: the answer must satisfy every clue. The runner also stops a model that solves none of the 5x5 puzzles with structured output, and one whose early runs mostly report zero reasoning tokens, since both point at the harness rather than the model.

After benchmarking, export results for the visualizer:

bash
bun run export

This writes visualizer/app/results.json (aggregates) and visualizer/public/results-raw.json (every run, including prompts and outputs).

Other scripts:

  • bun test - parser, grader, database-policy and puzzle checks
  • bun run typecheck - TypeScript check
  • bun run regrade - read-only comparison of stored grades against the current grader

4. Viewing Results

bash
cd visualizerbun installbun run dev

Then open http://localhost:3000 to view the interactive dashboard.

Agent Access

nonobench.com exposes the benchmark data to agents, with no authentication:

  • REST API under /api/v1 (leaderboard, models, puzzles, a solution checker, individual runs). The spec is at /api/openapi.json, and /.well-known/api-catalog (RFC 9727) points to it.

  • MCP server at /mcp (stateless Streamable HTTP, MCP 2026-07-28 with 2025 client compatibility), described by /.well-known/mcp/server-card.json. Add it to a client with claude mcp add --transport http nonobench https://www.nonobench.com/mcp. Browser requests may use the site's origins or HTTP localhost/127.0.0.1 origins. It's listed in the Claude Connectors Directory, the MCP Registry (io.github.mauricekleine/nonobench), Smithery, Glama and mcpservers.org.

    [Nonobench MCP connector] [Listed on mcpservers.org]

  • WebMCP tools registered in the browser via navigator.modelContext.

  • Markdown: / and /puzzles return markdown when requested with Accept: text/markdown. /llms.txt gives an overview.

  • Discovery: robots.txt (with Content Signals), sitemap.xml, Link headers on the homepage, an agent skill at /.well-known/agent-skills/index.json, and an ARD manifest at /.well-known/ai-catalog.json.

All of it is read from the same exported files as the dashboard (visualizer/app/results.json and visualizer/public/results-raw.json), so bun run export updates it too.

Grading

Each model receives the same system prompt and the puzzle's row and column clues. Standard answers are the grid as one string of 1s and 0s. Hard mode answers are one row per line, because at 400 cells most models miscount a single string (see LEARNINGS.md). An answer is correct when it satisfies every row and column clue.

Ten of the 30 puzzles (one 5x5, four 10x10, five 15x15) have more than one valid solution, so answers are checked against the clues rather than compared with the stored solution. Correctness is derived from the stored raw outputs at export time; the database is never rewritten. The puzzle test suite pins which puzzles are ambiguous, and any new puzzle must have a unique solution.

Puzzle Data

The core tier has 30 puzzles (10 each of 5x5, 10x10 and 15x15), defined in visualizer/components/puzzles/ and shared by the runner and the dashboard. They were sourced from nono-dataset. Hard mode has 10 generated 20x20 puzzles. A puzzle's ID is a hash of its solution, so changing a puzzle's solution creates a new puzzle.

Tiers and generation

Default benchmark runs cover the three core sizes (Standard). Use --sizes 20x20 with a model selection to run Hard mode; comma-separated sizes also work. The runner's plan reports missing 20x20 work separately. Headline overall accuracy and best-variant selection use Standard runs only; Hard mode has its own results. Hard-mode requests get a 128,000-token answer budget, capped at the endpoint's maximum (bench/max-output-tokens.json).

From bench/, bun run generate-puzzles recreates the 20x20 set with a fixed seed. It fills grids at random (no pictures, so a model can't guess the image), keeps grids with at least three blocks per line and little mirror symmetry, and checks uniqueness with an exact solver (NONOGRAM_SOLVER). The set mixes five puzzles that row-and-column propagation solves with five where it stalls with 20–200 cells left. bun test verifies their clues, uniqueness flags and line solvability; the original ambiguity list remains pinned.

Configuration

Edit bench/constants.ts to configure:

  • MODELS - Array of model configurations (OpenRouter model ID, display name, reasoning settings)
  • MAX_PARALLEL_RUNS_PER_MODEL - Concurrent puzzle runs per model (default: 10)
  • REQUEST_TIMEOUT_MS - Per-request timeout; a timed-out request is stored as a failed run (default: 30 minutes)

NONOBENCH_DB, NONOBENCH_RESULTS_JSON and NONOBENCH_RESULTS_RAW_JSON override the database and export paths, which is useful for testing against a copy.

Tech Stack

Benchmark Runner

  • Bun - JavaScript runtime and SQLite
  • AI SDK - Unified LLM interface
  • OpenRouter - LLM API gateway
  • TypeScript

Visualizer

Contributing

Contributions are welcome! Feel free to:

  • Add support for new LLM models
  • Improve the benchmark methodology
  • Enhance the visualization dashboard

License

MIT

來源:README.md,提交 4610b89

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v1.1.0最新Sep 30, 2026