Nonobench

io.github.mauricekleinev1.1.0Updated Sep 30, 2026

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

VerifiedStreamable HTTPWeb executableAI & MLData & Analytics

Overview

AI-generated overview

Read-only access to the Nonobench benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

What it does
Nonobench is an open-source benchmark measuring LLM reasoning on nonogram (Picross) puzzles. The hosted MCP endpoint exposes the published benchmark data: leaderboard, models, puzzles, individual runs, and a solution checker, all read from the same exported result files as the dashboard. It is stateless Streamable HTTP and requires no authentication.
When to use it
Use it when you want an assistant to look up nonogram benchmark results, compare model performance on these puzzles, inspect individual runs, or check a candidate solution against the clues. It is a data-reading endpoint, not a tool for running new benchmarks.
Requirements
A remote MCP client pointed at the hosted endpoint; no account, API key, or header is declared. Running the benchmark itself is separate and needs the repository, the Bun runtime, Node.js for the visualizer, and an OpenRouter API key.
Before you install
The endpoint is read-only and unauthenticated, so nothing sensitive is sent. Note that the exported raw results include prompts and model outputs, and that running the benchmark locally spends money through OpenRouter; the OPENROUTER_API_KEY is the credential involved, and the key's own limit is the hard spending ceiling.

Installation

In SourceWeft

  1. Open Nonobench in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.

Other MCP clients

Add this to your client's mcpServers config.

{
  "mcpServers": {
    "nonobench": {
      "type": "http",
      "url": "https://www.nonobench.com/mcp"
    }
  }
}

README

Nonobench

A benchmark suite for evaluating LLM reasoning capabilities on Nonogram (Picross) puzzle solving across different grid sizes. Results are published at nonobench.com.

Built by Maurice Kleine.

What is a Nonogram?

Nonograms (also known as Picross, Griddlers, or Paint by Numbers) are logic puzzles where you fill in cells on a grid based on numeric clues for each row and column. The clues indicate consecutive groups of filled cells, separated by at least one empty cell. Solving these puzzles requires logical deduction and constraint satisfaction - making them an excellent test of LLM reasoning abilities.

Project Structure

nonobench/├── bench/          # Benchmark runner, results database and exporter└── visualizer/     # Next.js dashboard (nonobench.com), also home of the puzzle set

Prerequisites

Quick Start

1. Clone the Repository

bash
git clone https://github.com/mauricekleine/nonobench.gitcd nonobench

2. Set Up Environment Variables

Create a .env file in the bench/ directory (see bench/.env.example):

bash
OPENROUTER_API_KEY=your_openrouter_api_key_here

This is the only variable needed. The visualizer builds without any.

3. Running Benchmarks

bash
cd benchbun installbun run bench                       # prints the plan and exits, no API callsbun run bench --model <name>        # run one model (repeat --model for more)bun run bench --all-missing         # run every configured model with missing workbun run bench --model <name> --sizes 20x20  # opt in to the extended tier

Runs are incremental and append-only: a model/puzzle pair that already has a successful result is never run again, and the database refuses to overwrite it. Failed runs are retried on the next invocation. Results are stored in bench/results.db (SQLite).

Useful flags:

  • --max-cost <usd> stops launching new puzzles once this session's spend reaches the amount. Requests already in flight still finish, so a session can overshoot by up to --parallel requests per selected model. Use the OpenRouter key's own limit as the hard ceiling.
  • --parallel <n> sets concurrent requests per model (default 10); lower it for rate-limited providers.
  • --limit <n> runs only the first n puzzles of each size, for pilots against a scratch database (NONOBENCH_DB=/tmp/copy.db).

Output modes

New runs ask for the answer as strict structured output (a JSON schema, only routed to endpoints that enforce it), so models cannot wrap the grid in prose. For a few models the schema-enforcing endpoints measurably hurt answers; those run in text mode instead (outputMode: "text" in bench/constants.ts, chosen by a 5x5 A/B with the benchmark prompt). Every run records its mode, and grading is identical for both: the answer must satisfy every clue. The runner also stops a model that solves none of the 5x5 puzzles with structured output, and one whose early runs mostly report zero reasoning tokens, since both point at the harness rather than the model.

After benchmarking, export results for the visualizer:

bash
bun run export

This writes visualizer/app/results.json (aggregates) and visualizer/public/results-raw.json (every run, including prompts and outputs).

Other scripts:

  • bun test - parser, grader, database-policy and puzzle checks
  • bun run typecheck - TypeScript check
  • bun run regrade - read-only comparison of stored grades against the current grader

4. Viewing Results

bash
cd visualizerbun installbun run dev

Then open http://localhost:3000 to view the interactive dashboard.

Agent Access

nonobench.com exposes the benchmark data to agents, with no authentication:

  • REST API under /api/v1 (leaderboard, models, puzzles, a solution checker, individual runs). The spec is at /api/openapi.json, and /.well-known/api-catalog (RFC 9727) points to it.

  • MCP server at /mcp (stateless Streamable HTTP, MCP 2026-07-28 with 2025 client compatibility), described by /.well-known/mcp/server-card.json. Add it to a client with claude mcp add --transport http nonobench https://www.nonobench.com/mcp. Browser requests may use the site's origins or HTTP localhost/127.0.0.1 origins. It's listed in the Claude Connectors Directory, the MCP Registry (io.github.mauricekleine/nonobench), Smithery, Glama and mcpservers.org.

    [Nonobench MCP connector] [Listed on mcpservers.org]

  • WebMCP tools registered in the browser via navigator.modelContext.

  • Markdown: / and /puzzles return markdown when requested with Accept: text/markdown. /llms.txt gives an overview.

  • Discovery: robots.txt (with Content Signals), sitemap.xml, Link headers on the homepage, an agent skill at /.well-known/agent-skills/index.json, and an ARD manifest at /.well-known/ai-catalog.json.

All of it is read from the same exported files as the dashboard (visualizer/app/results.json and visualizer/public/results-raw.json), so bun run export updates it too.

Grading

Each model receives the same system prompt and the puzzle's row and column clues. Standard answers are the grid as one string of 1s and 0s. Hard mode answers are one row per line, because at 400 cells most models miscount a single string (see LEARNINGS.md). An answer is correct when it satisfies every row and column clue.

Ten of the 30 puzzles (one 5x5, four 10x10, five 15x15) have more than one valid solution, so answers are checked against the clues rather than compared with the stored solution. Correctness is derived from the stored raw outputs at export time; the database is never rewritten. The puzzle test suite pins which puzzles are ambiguous, and any new puzzle must have a unique solution.

Puzzle Data

The core tier has 30 puzzles (10 each of 5x5, 10x10 and 15x15), defined in visualizer/components/puzzles/ and shared by the runner and the dashboard. They were sourced from nono-dataset. Hard mode has 10 generated 20x20 puzzles. A puzzle's ID is a hash of its solution, so changing a puzzle's solution creates a new puzzle.

Tiers and generation

Default benchmark runs cover the three core sizes (Standard). Use --sizes 20x20 with a model selection to run Hard mode; comma-separated sizes also work. The runner's plan reports missing 20x20 work separately. Headline overall accuracy and best-variant selection use Standard runs only; Hard mode has its own results. Hard-mode requests get a 128,000-token answer budget, capped at the endpoint's maximum (bench/max-output-tokens.json).

From bench/, bun run generate-puzzles recreates the 20x20 set with a fixed seed. It fills grids at random (no pictures, so a model can't guess the image), keeps grids with at least three blocks per line and little mirror symmetry, and checks uniqueness with an exact solver (NONOGRAM_SOLVER). The set mixes five puzzles that row-and-column propagation solves with five where it stalls with 20–200 cells left. bun test verifies their clues, uniqueness flags and line solvability; the original ambiguity list remains pinned.

Configuration

Edit bench/constants.ts to configure:

  • MODELS - Array of model configurations (OpenRouter model ID, display name, reasoning settings)
  • MAX_PARALLEL_RUNS_PER_MODEL - Concurrent puzzle runs per model (default: 10)
  • REQUEST_TIMEOUT_MS - Per-request timeout; a timed-out request is stored as a failed run (default: 30 minutes)

NONOBENCH_DB, NONOBENCH_RESULTS_JSON and NONOBENCH_RESULTS_RAW_JSON override the database and export paths, which is useful for testing against a copy.

Tech Stack

Benchmark Runner

  • Bun - JavaScript runtime and SQLite
  • AI SDK - Unified LLM interface
  • OpenRouter - LLM API gateway
  • TypeScript

Visualizer

Contributing

Contributions are welcome! Feel free to:

  • Add support for new LLM models
  • Improve the benchmark methodology
  • Enhance the visualization dashboard

License

MIT

Source: README.md at commit 4610b89

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v1.1.0LatestSep 30, 2026