Nonobench

io.github.mauricekleinev1.1.0更新于 Sep 30, 2026

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

已验证Streamable HTTP可网页运行AI & MLData & Analytics

概览

AI 生成的概览

只读访问 Nonobench 基准数据,了解大语言模型在 5x5 至 20x20 数织谜题上的表现。

功能
Nonobench 是一个开源基准,用于评估大语言模型在数织(Picross)谜题上的推理能力。托管的 MCP 端点提供已发布的基准数据:排行榜、模型、谜题、单次运行记录以及解答校验器,数据与仪表盘来自同一批导出文件。该端点是无状态 Streamable HTTP,无需身份验证。
适用场景
当你希望助手查询数织基准结果、比较各模型在这些谜题上的表现、查看单次运行记录,或根据线索校验某个解答时,可以使用它。它是读取数据的端点,不能用来运行新的基准测试。
运行要求
需要一个指向托管端点的远程 MCP 客户端;未声明账号、API 密钥或请求头。运行基准测试本身是另一回事,需要仓库、Bun 运行时、用于可视化界面的 Node.js,以及 OpenRouter API 密钥。
安装前请注意
该端点只读且无需认证,因此不会发送敏感信息。注意导出的原始结果包含提示词和模型输出;在本地运行基准测试会通过 OpenRouter 产生费用,涉及的凭据是 OPENROUTER_API_KEY,密钥自身的额度上限才是支出的硬性上限。

安装

在 SourceWeft 中

  1. 打开 控制台中的 Nonobench,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Web executable,通过 Streamable HTTP。 远程服务在工作区中配置后即可从网页运行时运行。

其他 MCP 客户端

把它添加到你客户端的 mcpServers 配置中。

{
  "mcpServers": {
    "nonobench": {
      "type": "http",
      "url": "https://www.nonobench.com/mcp"
    }
  }
}

README

Nonobench

A benchmark suite for evaluating LLM reasoning capabilities on Nonogram (Picross) puzzle solving across different grid sizes. Results are published at nonobench.com.

Built by Maurice Kleine.

What is a Nonogram?

Nonograms (also known as Picross, Griddlers, or Paint by Numbers) are logic puzzles where you fill in cells on a grid based on numeric clues for each row and column. The clues indicate consecutive groups of filled cells, separated by at least one empty cell. Solving these puzzles requires logical deduction and constraint satisfaction - making them an excellent test of LLM reasoning abilities.

Project Structure

nonobench/├── bench/          # Benchmark runner, results database and exporter└── visualizer/     # Next.js dashboard (nonobench.com), also home of the puzzle set

Prerequisites

Quick Start

1. Clone the Repository

bash
git clone https://github.com/mauricekleine/nonobench.gitcd nonobench

2. Set Up Environment Variables

Create a .env file in the bench/ directory (see bench/.env.example):

bash
OPENROUTER_API_KEY=your_openrouter_api_key_here

This is the only variable needed. The visualizer builds without any.

3. Running Benchmarks

bash
cd benchbun installbun run bench                       # prints the plan and exits, no API callsbun run bench --model <name>        # run one model (repeat --model for more)bun run bench --all-missing         # run every configured model with missing workbun run bench --model <name> --sizes 20x20  # opt in to the extended tier

Runs are incremental and append-only: a model/puzzle pair that already has a successful result is never run again, and the database refuses to overwrite it. Failed runs are retried on the next invocation. Results are stored in bench/results.db (SQLite).

Useful flags:

  • --max-cost <usd> stops launching new puzzles once this session's spend reaches the amount. Requests already in flight still finish, so a session can overshoot by up to --parallel requests per selected model. Use the OpenRouter key's own limit as the hard ceiling.
  • --parallel <n> sets concurrent requests per model (default 10); lower it for rate-limited providers.
  • --limit <n> runs only the first n puzzles of each size, for pilots against a scratch database (NONOBENCH_DB=/tmp/copy.db).

Output modes

New runs ask for the answer as strict structured output (a JSON schema, only routed to endpoints that enforce it), so models cannot wrap the grid in prose. For a few models the schema-enforcing endpoints measurably hurt answers; those run in text mode instead (outputMode: "text" in bench/constants.ts, chosen by a 5x5 A/B with the benchmark prompt). Every run records its mode, and grading is identical for both: the answer must satisfy every clue. The runner also stops a model that solves none of the 5x5 puzzles with structured output, and one whose early runs mostly report zero reasoning tokens, since both point at the harness rather than the model.

After benchmarking, export results for the visualizer:

bash
bun run export

This writes visualizer/app/results.json (aggregates) and visualizer/public/results-raw.json (every run, including prompts and outputs).

Other scripts:

  • bun test - parser, grader, database-policy and puzzle checks
  • bun run typecheck - TypeScript check
  • bun run regrade - read-only comparison of stored grades against the current grader

4. Viewing Results

bash
cd visualizerbun installbun run dev

Then open http://localhost:3000 to view the interactive dashboard.

Agent Access

nonobench.com exposes the benchmark data to agents, with no authentication:

  • REST API under /api/v1 (leaderboard, models, puzzles, a solution checker, individual runs). The spec is at /api/openapi.json, and /.well-known/api-catalog (RFC 9727) points to it.

  • MCP server at /mcp (stateless Streamable HTTP, MCP 2026-07-28 with 2025 client compatibility), described by /.well-known/mcp/server-card.json. Add it to a client with claude mcp add --transport http nonobench https://www.nonobench.com/mcp. Browser requests may use the site's origins or HTTP localhost/127.0.0.1 origins. It's listed in the Claude Connectors Directory, the MCP Registry (io.github.mauricekleine/nonobench), Smithery, Glama and mcpservers.org.

    [Nonobench MCP connector] [Listed on mcpservers.org]

  • WebMCP tools registered in the browser via navigator.modelContext.

  • Markdown: / and /puzzles return markdown when requested with Accept: text/markdown. /llms.txt gives an overview.

  • Discovery: robots.txt (with Content Signals), sitemap.xml, Link headers on the homepage, an agent skill at /.well-known/agent-skills/index.json, and an ARD manifest at /.well-known/ai-catalog.json.

All of it is read from the same exported files as the dashboard (visualizer/app/results.json and visualizer/public/results-raw.json), so bun run export updates it too.

Grading

Each model receives the same system prompt and the puzzle's row and column clues. Standard answers are the grid as one string of 1s and 0s. Hard mode answers are one row per line, because at 400 cells most models miscount a single string (see LEARNINGS.md). An answer is correct when it satisfies every row and column clue.

Ten of the 30 puzzles (one 5x5, four 10x10, five 15x15) have more than one valid solution, so answers are checked against the clues rather than compared with the stored solution. Correctness is derived from the stored raw outputs at export time; the database is never rewritten. The puzzle test suite pins which puzzles are ambiguous, and any new puzzle must have a unique solution.

Puzzle Data

The core tier has 30 puzzles (10 each of 5x5, 10x10 and 15x15), defined in visualizer/components/puzzles/ and shared by the runner and the dashboard. They were sourced from nono-dataset. Hard mode has 10 generated 20x20 puzzles. A puzzle's ID is a hash of its solution, so changing a puzzle's solution creates a new puzzle.

Tiers and generation

Default benchmark runs cover the three core sizes (Standard). Use --sizes 20x20 with a model selection to run Hard mode; comma-separated sizes also work. The runner's plan reports missing 20x20 work separately. Headline overall accuracy and best-variant selection use Standard runs only; Hard mode has its own results. Hard-mode requests get a 128,000-token answer budget, capped at the endpoint's maximum (bench/max-output-tokens.json).

From bench/, bun run generate-puzzles recreates the 20x20 set with a fixed seed. It fills grids at random (no pictures, so a model can't guess the image), keeps grids with at least three blocks per line and little mirror symmetry, and checks uniqueness with an exact solver (NONOGRAM_SOLVER). The set mixes five puzzles that row-and-column propagation solves with five where it stalls with 20–200 cells left. bun test verifies their clues, uniqueness flags and line solvability; the original ambiguity list remains pinned.

Configuration

Edit bench/constants.ts to configure:

  • MODELS - Array of model configurations (OpenRouter model ID, display name, reasoning settings)
  • MAX_PARALLEL_RUNS_PER_MODEL - Concurrent puzzle runs per model (default: 10)
  • REQUEST_TIMEOUT_MS - Per-request timeout; a timed-out request is stored as a failed run (default: 30 minutes)

NONOBENCH_DB, NONOBENCH_RESULTS_JSON and NONOBENCH_RESULTS_RAW_JSON override the database and export paths, which is useful for testing against a copy.

Tech Stack

Benchmark Runner

  • Bun - JavaScript runtime and SQLite
  • AI SDK - Unified LLM interface
  • OpenRouter - LLM API gateway
  • TypeScript

Visualizer

Contributing

Contributions are welcome! Feel free to:

  • Add support for new LLM models
  • Improve the benchmark methodology
  • Enhance the visualization dashboard

License

MIT

来源:README.md,提交 4610b89

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v1.1.0最新Sep 30, 2026