
Design Pattern MCP
io.github.olkv1.0.1更新於 Oct 10, 2026
MCP server that provides design-pattern selection expertise to AI coding agents
概覽
讓程式開發助理從 438 筆設計模式目錄中挑選模式,並為所描述的系統產生結構化設計文件。
- 功能
- 提供多項工具,接收自由文字的系统描述並回傳設計結果:包含主要模式與輔助模式的概觀、含參與者、互動、實作步驟與程式碼草稿的應用方案、品質屬性、測試策略與風險。工具包括 select_design_patterns、analyze_design、generate_design、evaluate_design、list_design_patterns 與 get_design_pattern,以及背景工作工具 submit_design_job、get_design_status 與 cancel_design。此外還提供四個由使用者主動呼叫的工作流程提示。
- 適用情境
- 當助理需要為所描述的系統推薦架構或設計模式、比較候選模式,或依品質面向評估既有設計時使用。它適合設計探索與審查,而非快速查詢,因為主要流程每次呼叫需執行數分鐘。
- 執行需求
- 以 Docker 映像或透過 uv 安裝的 Python 3.12+ 套件在本機執行,支援 stdio 或 streamable-http。需要產生器 LLM 的 API 金鑰(DESIGN_PATTERN_GENERATOR_API_KEY)以及供應商、模型與基礎 URL 設定,並需要可連線的 TEI 嵌入與重排端點(DESIGN_PATTERN_EMBEDDER_BASE_URL、DESIGN_PATTERN_RERANKER_BASE_URL)。Docker 安裝需 Docker v2.20+、約 8 GB 可用空間與 linux/amd64。另有僅詞法的離線模式。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Design Pattern MCP,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
design-pattern-mcp
An MCP (Model Context Protocol) server that provides design-pattern selection expertise to AI coding agents. Given a free-text description of a system, it analyses the problem, retrieves matching design patterns (from 438 built-in patterns, each realized in 13 concrete languages plus a generic binding), generates a concrete design with a primary pattern, supporting patterns, participants, interactions, implementation steps, and language-bound variants, and evaluates it against eight design-level quality axes (testability, extensibility, maintainability, simplicity, performance, reliability, security, scalability). The pipeline (analyze → weights ∥ retrieval → generate → evaluate → bounded refine) and its wire contract are specified in docs/implementation-guide.md.
Table of Contents
- ⚡ Quickstart
- Install via AI Agent
- 🔌 Connect Your Agent
- 🧑🏫 SKILL for AI Agents
- 🧪 Use the Tools
- 🛠️ Tools at a Glance
- 💬 Prompts
- 📖 Pattern Catalog
- Working Together with architecture-pattern-mcp
- Install Alternatives
- Configuration
- Structured Reasoning
- Long-running tools & timeouts
- Benchmarks & Verification
- Troubleshooting
- Building & Development
- Publishing
- systemd Service (Linux)
- License
⚡ Quickstart
Server starts on streamable-http at http://localhost:8062/mcp (dev compose maps host MCP_HOST_PORT:-8062 → container MCP_CONTAINER_PORT:-8052). make client targets http://localhost:8062/mcp; override with DESIGN_PATTERN_CLIENT_URL.
Install via AI Agent
To have an AI coding agent install the Docker stack from prebuilt images (no local build), point it at INSTALL.md — the non-interactive agent runbook (copy-paste commands, expected outputs, verification probes).
Pick one path in INSTALL.md:
- Path A — Compose stack (
:8062): evaluating / developing, no sudo needed. Own TEI sidecar copies. - Path B — systemd stack (
:8062): persistent host service, starts at boot, needssudo. Uses the sharedpattern-tei-infraTEI stack.
The agent needs a repo clone, Docker ≥ v2.20 with group membership, ~8 GB free, a MINIMAXAI_API_KEY (or another provider per INSTALL.md#switching-the-generator-llm), and linux/amd64 (on arm64 it builds locally via make docker-build-all). Both paths default to host port 8062 — set MCP_HOST_PORT in one env file to run both at once. After install, continue with Connect Your Agent below.
🔌 Connect Your Agent
Claude Code
Stdio mode needs the TEI endpoints reachable (default http://design-pattern-tei-embed:8080 and the reranker URL from config/config.json) — either run the compose stack's sidecars or set DESIGN_PATTERN_EMBEDDER_BASE_URL / DESIGN_PATTERN_RERANKER_BASE_URL to your own. For a no-network smoke run, start with --offline (deterministic lexical-only retrieval instead of TEI).
OpenCode
OpenCode uses HTTP transport. Start the server first, then configure opencode:
Note: the generator API key is read by the server from its config file (
config/config.json,{env:...}placeholders), not from opencode's environment.
Codex CLI
Add to ~/.codex/config.toml:
Or via CLI:
🧑🏫 SKILL for AI Agents
The SKILL in skills/design-pattern-mcp/ is written for Oh My Pi (OMP), where this server's tools are reached as xd://mcp__design_pattern_local_* devices and every MCP request is bounded by a deadline (OMP_MCP_TIMEOUT_MS → per-server timeout → 30 s). It teaches the agent which entry point fits that deadline, how to read the results, and the full workflow recipes.
Install for OMP — copy the skill directory into OMP's user skills root:
The skill then tells the agent:
- Which entry point fits OMP's 30 s deadline: the async job trio by default;
select_design_patterns(full pipeline) takes 2–10 min - That
job_idis the only durable handle — persist it (SQLite job store,DESIGN_PATTERN_JOBS_DB) - How to phrase
description(component names + responsibilities + constraints), andlanguage,paradigm,focusas separate structured arguments - How to interpret
overview.score(0–100),findings[], and per-part selections
The tool schemas in references/ are client-agnostic; only the deadline/device guidance is OMP-specific.
🧪 Use the Tools
All pipeline tools take a free-text description as the primary argument. The examples below show the exact tool call shape so you can use them in any MCP client or API consumer.
Design your first selection
In Claude Code (or any MCP client), paste the natural-language instruction:
Your agent calls select_design_patterns internally. The server returns the DesignResult: overview (primary pattern, supporting patterns, principles, score, reasoning), applications[] (participants, interactions, implementation_steps, code_sketch, language_binding with the bound variant), interactions[], quality_attributes, test_strategy, and risks[].
Or call tools directly from your agent:
Async job pattern: submit_design_job + get_design_status
For clients with short request timeouts. submit_design_job returns a job_id immediately; poll get_design_status until done:
In Python (via the MCP client API — see examples/design_pattern_client_async.py):
See examples/design_pattern_client_async.py for the complete runnable example. Run it with:
🛠️ Tools at a Glance
Description, language, paradigm, focus are structured parameters — pass them as separate tool arguments, not embedded in the description text. Variant selection is signal-based: there is no catalog-wide variant_id enum (variant ids are per-(pattern, language) and change with catalog edits; valid ids are discoverable via the design-patterns://variants resource and previous runs' bound_variant).
💬 Prompts
This server also exposes four user-invoked workflow prompts (slash commands in MCP clients). Unlike tools, the LLM does not autonomously invoke prompts — the user selects one and fills in its arguments. Each prompt encodes a tested tool-orchestration recipe, and none hardcode pattern names (category guidance quotes the live catalog's own structure).
* = required argument
Tool-only clients
Clients that only support the tools protocol (no native prompts/list or prompts/get) can access all four workflow prompts via the generated list_prompts and get_prompt tools, which route through the server's middleware chain exactly as native prompt calls do.
📖 Pattern Catalog
Via MCP tools (recommended — works in all clients)
Valid category values (the nine decision groups, catalog sections 1–9): domain-modeling, construction-composition, behavior-selection, interaction-boundary, workflow-messaging, persistence-state, concurrency-resilience, security-verification, specialized-components.
Via MCP resources
Groups every pattern's language realizations by language, then pattern — the realization vocabulary for language/paradigm/focus hints and variant selection.
What one pattern record contains
The loader (src/patterns/loader.py) joins each record from 14 files: one generic entry plus one per concrete language (13 languages: C, C++, C#, Go, Java, JavaScript, Kotlin, Python, Ruby, Rust, Scala, Swift, TypeScript). Each language binding carries support (native / idiomatic / discouraged), idiom_name, implementation_notes, an optional code_sketch, pitfalls, and — for discouraged support — a required alternative. Language bindings can declare variants[] (per-language realizations, 4,924 across the catalog) with a default_variant. Each record also declares related_patterns (loader-resolvable canonical names only) and related_patterns_notes (per-link prose distinctions).
Alias resolution is built in: the loader turns 841 raw alias entries → 827 distinct after case-folding → 816 indexed alias mappings, dropping ambiguous aliases (an alias that is itself another record's name or claimed by two records; the ambiguous_aliases health field lists them), so get_design_pattern(name="circuit-breaker") and its alias spellings resolve to the same record. Startup fails fast when the catalog's enums drift from docs/pattern-schema.json (assert_enums_in_sync).
Working Together with architecture-pattern-mcp
This server answers how each unit is coded (design patterns); the sibling architecture-pattern-mcp answers which system shape to build (architecture). An architecture pattern structures the whole system — its components, their relationships, and its API/data/event contracts (e.g. event-driven, pipe-and-filter, microservices, chosen from the sibling's 40-pattern catalog). A design pattern solves one recurring implementation problem inside a single component — its participants, interactions, and language-bound code (e.g. observer, process-manager, circuit-breaker, chosen from the 438-pattern catalog above with 13 language bindings each).
Recommended order — macro first, micro second:
Concrete handoff on the IoT example: the sibling emits the pipe-and-filter architecture (Kafka source → JSON parser filter → geolocation enricher → InfluxDB/S3 sinks, with at-least-once delivery and backpressure in the contracts). The geolocation enricher's cache-miss path then comes here via select_design_patterns ("cache-aside enrichment with Redis fallback, Python") — e.g. a circuit-breaker + retry design with participants, implementation_steps, and a Python code_sketch — which the architecture-level component contract references but never implements.
Run both servers side by side (dev host ports :8060 architecture, :8062 design-pattern; systemd :8050 / :8062). They are independent — either order works for exploration — but architecture-first keeps component boundaries stable while this server fills in the internals.
Install Alternatives
Docker (manual)
First docker-build-all generates the FAISS dense cache (data/dense-cache/) via the TEI embed sidecar and bakes it into the server image; make docker-build-nocache skips that (first startup then embeds the corpus, ~19 min).
Local Development (uv)
Prerequisites: Python 3.12+, uv
Or use the installed console script (after make install): design-pattern-mcp.
The TEI sidecars are required for full retrieval quality: a dense leg (Qwen3-Embedding-0.6B via the embed sidecar) and a BM25 leg are fused at startup (service_healthy ordering guarantees the embedder is up before the server accepts requests — a failed index warm-up refuses server start). The reranker sidecar (gte-reranker-modernbert-base) reranks candidate patterns. Without reachable sidecars, start with --offline (deterministic lexical-only retrieval, no network).
TEI startup: both sidecar images (
docker/Dockerfile.tei-embed,docker/Dockerfile.tei-rerank) ship a self-built TEI router (v1.9.4 + unmerged PR #884) withWARMUP_TOKENS=0by default — this skips the synthetic warmup pass (measured ~131 s embed / ~78 s rerank) so containers become ready after weight load only. Override per sidecar viaTEI_WARMUP_TOKENS/TEI_RERANK_WARMUP_TOKENS(positive N warms with N tokens instead; values >MAX_BATCH_TOKENSare rejected at startup).
GPU TEI: the shipped sidecars are CPU images (cpu-1.9). On a host with a TEI-CUDA-supported NVIDIA GPU (compute capability ≥ 7.5) plus the NVIDIA container toolkit, make docker-build-tei-embed-gpu builds the embed sidecar on the CUDA base (ghcr.io/huggingface/text-embeddings-inference:cuda-1.9) and make docker-up-gpu starts the dev stack with docker/docker-compose.gpu.yml (reserves 1 GPU). The override covers the embed sidecar only — the reranker stays on CPU. On GPU a small positive WARMUP_TOKENS warms kernels at startup instead of skipping the pass.
Already-running embedder/reranker: when embedding and reranking endpoints already run on other servers, point the MCP server at them via DESIGN_PATTERN_EMBEDDER_BASE_URL (OpenAI-compatible route, http://<host>:8080/v1) plus DESIGN_PATTERN_RERANKER_BASE_URL (http://<host>:8080) and skip the sidecars. The embedder forwards DESIGN_PATTERN_EMBEDDER_MODEL (default openai//data/embedding-model) and DESIGN_PATTERN_EMBEDDER_API_KEY through LiteLLM — provider/model syntax is documented in the LiteLLM embedding docs. The reranker is not a LiteLLM client — it POSTs {base_url}/rerank directly (src/patterns/safe_tei_rerank.py), so DESIGN_PATTERN_RERANKER_MODEL (default Alibaba-NLP/gte-reranker-modernbert-base) is a TEI model name, not a LiteLLM model string. A remote embedder serving a different model invalidates the baked dense cache — regenerate it with make embed-cache so vectors match the serving model.
Configuration
config.json
The server reads config/config.json relative to the repo (override with --config-path or DESIGN_PATTERN_CONFIG_PATH). Every value is a {env:VAR:-default} placeholder that expands at load time, so the shipped file doubles as the env-var reference. The schema (abridged):
Generator LLM (LlamaIndex LiteLLM)
The generator LLM is accessed through the LlamaIndex LiteLLM integration. All provider settings follow LiteLLM's model syntax: <provider>/<model>. The dev stack and benchmark harness default to MiniMax:
Provider list and syntax: LiteLLM Providers documentation. temperature, top_p, top_k, timeout_seconds, and the circuit-breaker knobs map to the corresponding LiteLLM/provider parameters.
Key environment variables
CLI flags
Structured Reasoning
Before each prompted LLM phase call (analyze / weights / generate / evaluate), the server runs a bounded ThoughtGenerator loop: each step is authored with the generator LLM (one completion per step) and submitted to the shannonthinking and/or code-reasoning MCP servers — structured thinking scratchpads that validate, number, and record each step. The resulting trace is injected into the phase prompt as a <reasoning_context> block. Contract: each thought = 1 LLM completion + 1 MCP tool call, capped by REASONING_MAX_TOTAL_STEPS (default 8). The tools list per phase defaults to both scratchpads; startup health-checks exactly the tools the enabled phases use.
Key properties:
- Embedded in Docker — the image bakes both npm packages into the image; the runtime invokes them directly via
node(no network, no npx). - Auto-fallback to npx outside Docker — missing embedded entry points fall back to
npx -y <pkg>(first call downloads). - Silent per-call degradation — any spawn/timeout/tool failure in an enabled phase logs a WARNING and the phase proceeds with a degraded in-prompt thinking scaffold; a reasoning outage never fails a tool call. A phase disabled via
REASONING_<PHASE>_ENABLED=falseinstead contributes an empty reasoning context — no generator call, no subprocess, no cache entry. - Loud startup, soft failure — the lifespan health-check reports reasoning health; set
REASONING_FAIL_FAST=trueto make a broken tool fail startup instead. - Trace caching — analyze and generate traces are computed once per design request and reused across design-loop attempts; weights and evaluate traces are per-run.
REASONING_MAX_INJECTED_CHARS— the rendered trace injected downstream can be capped (default uncapped) to cut input tokens on every downstream call.
Latency note: a full run issues 8 reasoning tool calls (2 per phase) as serial subprocesses; reasoning is a material share of end-to-end latency. Measured on the Stage-0 benchmark (MiniMax-M2.7 generator, 18 scenarios, NO_WARMUP=1): shipping config (all phases on) p50 288.4 s / p95 439.6 s; with REASONING_WEIGHTS_ENABLED=false REASONING_EVALUATE_ENABLED=false p50 253.0 s / p95 413.9 s — both arms passed the quality gate (overall hit ≥ baseline Wilson low). The per-phase toggles are the documented way to opt in.
Long-running tools & timeouts
select_design_patterns runs a multi-stage pipeline that takes minutes per call. This is inherent to the workload: the generator LLM processes the fused retrieval candidates, per-part quality weights, and the full output of every previous stage, and produces a large, strictly structured JSON document one token at a time. The design loop repeats generate → evaluate up to DESIGN_PATTERN_LOOP_MAX_TRIES times.
Four bounds govern a run — do not confuse them
The heartbeat defence (applied by default)
Every long-running tool emits progress notifications from a parallel coroutine every 30 seconds (DESIGN_PATTERN_TASKS_HEARTBEAT_INTERVAL_SECONDS). Clients whose idle timers reset on any received data stay alive for the full pipeline duration.
TS-SDK clients (Claude Desktop, Cursor) do not reset their timeout on progress notifications.
The async job trio (for timeout-constrained clients)
submit_design_job returns a job_id in milliseconds; the pipeline runs in a background task and the store is SQLite (DESIGN_PATTERN_JOBS_DB). Poll every 10–30 seconds. This is the fix that works for TS-SDK clients.
Bypassing client timeouts entirely: make client
examples/design_pattern_client.py is a direct Python HTTP client with no client-side idle timeout — one blocking request, full response regardless of duration:
Benchmarks & Verification
This repo treats selection quality and pipeline latency as measured properties, not claims.
Benchmark harness (tests/benchmark/, make benchmark-*): an 18-scenario corpus with holdout split; make benchmark-live NO_WARMUP=1 OUT=benchmarks/<arm> RUN_ID=<id> runs the full pipeline in-process with per-stage attribution (reasoning., llm., rerank, embed — calls, latency, chars/s, repair counters). make benchmark-selfcheck hand-verifies every metric formula. make benchmark-compare A=<baseline> B=<candidate> applies the preregistered rule — candidate wins iff holdout p95 e2e improves ≥ 15 % AND candidate overall acceptable-primary hit rate ≥ the baseline's overall Wilson lower bound — and exits 1 on candidate-loss. The e2e mode measures the deployed compose stack and is not comparable with the in-process mode.
Reference arms (committed under benchmarks/, all MiniMax-M2.7, 18 scenarios):
The code levers are shipped as opt-in env configuration, not defaults: the cumulative win claim passed the preregistered gate against base-fresh (holdout p95 410.9 → 297.5 s = 27.6 % ≥ 15 %, overall 13/18 ≥ floor), but the attribution comparison against cfg-reasoning recorded a candidate-loss (+12.5 % < 15 %), so defaults stay behavior-neutral. To opt in:
Verification layers (the repo's differentiator):
Every .fizz assertion must kill at least one spec-garden mutant or carry a # spec-explains: justification; a change to a decision module updates the matching .fizz model in the same PR. The normative pipeline spec is docs/implementation-guide.md (§-numbered, mirrored in German as docs/selektions-process-de.md); the wire contract of the tool result is docs/result-schema.json, the catalog record contract is docs/pattern-schema.json.
Troubleshooting
Server starts but tools are not visible
- Check the agent's MCP connection: Claude Code
/mcp, OpenCodeopencode mcp list, Codexcodex mcp list - Verify the server process started and the catalog loaded:
uv run design-pattern-mcp --healthprintsOKplus a JSON readiness report (records, language files, aliases) - Confirm the TEI embedder is healthy:
docker compose -f docker/docker-compose.yml ps(healthchecks with 120 s embed-model start period)
"Connection refused" or startup failure
Index warm-up refuses server start when the embedder is unreachable (fail-fast rather than breaking your first request):
Local non-Docker runs must point DESIGN_PATTERN_EMBEDDER_BASE_URL / DESIGN_PATTERN_RERANKER_BASE_URL at reachable endpoints — or start with --offline.
LLM provider errors (502 / 401)
- Confirm
DESIGN_PATTERN_GENERATOR_API_KEY(orMINIMAXAI_API_KEYon the dev stack) is set and not expired - Verify
DESIGN_PATTERN_GENERATOR_BASE_URLmatches your provider's endpoint - After 3 consecutive transport failures the circuit opens for 60 s (
..._CIRCUIT_BREAKER_*knobs); auth/configuration 4xx never count toward it
ERR_012 / ERR_014
ERR_012= cancelled (viacancel_designor a client disconnect at the next stage boundary)ERR_014= whole-run deadline exceeded (DESIGN_PATTERN_TASKS_DEADLINE_SECONDS); raise it explicitly for legitimately longer runs — the in-flight provider coroutine is cancelled, partial results are discarded
Reranker 429s
One full-text rerank request already saturates the CPU TEI sidecar's 16,384-token inference buffer (MAX_BATCH_TOKENS=8192 interacts with request packing; measured 2026-09-25): a cap-2 retrieve_many leg fails outright with HTTP 429 "Model is overloaded". Keep DESIGN_PATTERN_RETRIEVAL_MAX_CONCURRENT_PARTS=1 (shipped default); the server-side invariant also keeps dense_top_k + bm25_top_k ≤ the sidecar's MAX_CLIENT_BATCH_SIZE (48).
Pattern JSON files not loading
- Records are joined from 14 files per pattern (
pattern/<name>.generic.json+ 13<language>.json); the generic file is required - Validate a record against
docs/pattern-schema.json(scripts/validate_patterns.py) - Startup fails on enum drift from the schema (
assert_enums_in_sync) — that is intentional; fix the data, not the guard
Building & Development
Common make targets:
Development workflow:
Ground rules (from AGENTS.md): src/ is mypy-strict; schema changes change docs/implementation-guide.md + the schema + the code in the same commit; pipeline control-flow changes update the matching FizzBee model in the same PR.
Publishing
All three images are published to two registries simultaneously:
All three images share the same $(DOCKER_TAG) (the version from pyproject.toml). The TEI images are shared with the sibling architecture-pattern-mcp server and change rarely — blob deduplication keeps re-tagging unchanged TEI images cheap, and TEI pushes can be skipped when the sidecars are unchanged.
Bandwidth note: the TEI embedder image is ~5 GB (ONNX fp32 weights baked in). First push to each registry is ~5 GB upload. Subsequent pushes are incremental — only changed layers are transferred.
Prerequisites
Docker Hub — already authenticated locally (docker login).
GitHub Container Registry — requires a classic PAT with write:packages scope. 2FA is not an issue — PATs bypass it. After login the token is discarded; the credential persists in ~/.docker/config.json until you log out.
Publish (one-time setup + per-session)
On subsequent publishes repeat steps 1–3. If your PAT has expired, generate a new one at the link above. make docker-publish pushes only the MCP server image (plus the v$(DOCKER_TAG) git tag and the MCP Registry entry via mcp-publisher); make docker-publish-tei pushes only the two TEI sidecars.
First push — set packages public (GHCR only)
GHCR packages default to private. After the first make docker-publish-all, flip all three packages to public:
Set each to Public and save.
Partial failure recovery
If the push fails mid-way (e.g., GHCR auth was not configured), Docker Hub layers are already uploaded. After fixing auth, re-running make docker-publish-all is safe — each registry reports a cache hit for already-uploaded layers and completes the remaining push. For targeted retries, the TEI sidecars can be pushed with make docker-publish-tei.
systemd Service (Linux)
The server can run as a systemd service on any systemd-based Linux host. It starts the Docker Compose stack automatically at boot. Full agent-runbook install: INSTALL.md (Path B). systemd/ contains:
The production compose is the deployment variant of the dev file: pre-built images only, absolute paths, deployable to /etc/design-pattern-mcp/. The TEI sidecars live in the shared pattern-tei-infra stack; this stack joins its Docker network and reaches them by DNS name. Start the service only after the shared TEI infra is active (running).
Day-to-day: systemctl start|stop|restart design-pattern-mcp, logs via journalctl -u design-pattern-mcp -f. Full details in systemd/README.md.
License
MIT License. Copyright (c) 2026 Oliver Kowalke. See the license headers in the source files.
來源:README.md,提交 cb23214
工具
0版本歷史
1- v1.0.1最新Oct 10, 2026

