
CPU Performance Engineering
io.github.usamahzv0.1.1更新于 Oct 5, 2026
A CPU performance engineering brain: a vetted reading list, benchmarks and cited source passages
概览
让助手获得可检索的 CPU 性能工程阅读清单、基准测试和带引用的来源段落,用于诊断性能问题。
- 功能
- 把一份精选的 CPU 性能工程资料库作为可检索知识提供:条目及入选理由、基准测试、编辑记录和规则。ask、lookup、get_entry、reading_path、get_benchmark、check_evidence 等工具返回带页码和引用的段落。它还能解析粘贴的 perf stat、自顶向下、编译器提示和反汇编输出,计算指标并导向相关来源。首次运行会为所链接的论文和手册建立本地全文与语义索引。
- 适用场景
- 适合处理 CPU 性能问题、希望答案基于有引用的论文、手册和基准测试而非模型记忆的场景。可用于解读 perf stat 或自顶向下输出、编译器向量化提示、汇编代码,或为 NUMA 等主题规划阅读路径。
- 运行要求
- 以 stdio 在本地运行,客户端用 uvx cpu-perf(或 pip install cpu-perf)启动,因此需要 Python 和 uv。无需账号、API 密钥或环境变量。需要网络访问来构建所链接来源的本地资料库,可能耗时几分钟到一刻钟、占用数百兆字节;仅支持桌面客户端。
安装
在 SourceWeft 中
- 打开 控制台中的 CPU Performance Engineering,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
cpu-perf
An MCP server that gives any AI client the
whole CPU Performance Engineering
list as something it can search and reason over, instead of a page it has to
be pasted. Connect it to Claude Code, Codex, Claude Desktop, Cursor or VS
Code and use it for your own performance work: ask questions, paste perf stat or compiler
output, and the client's own model writes the answer from what the server
returns, with citations back to the sources.
It knows two things.
- The repository. Every entry and its reason, in reading order; the watchlist and what would promote each line; the editorial record behind the list (every candidate that was considered and left out, with the rule it failed; every performance number examined against the seven-field rule, with its verdict; the link-verification notes); the fourteen benchmarks with their claims, machines, results, analysis, code and raw output; and the house rules. This is bundled with the server and loads in a fraction of a second.
- The sources themselves. On first run the server reads every source the README links, plus the main document behind each link (the PDF behind an arXiv abstract, the manual behind a vendor landing page, a repository's README, a pull request's description), extracts the text with page numbers, and builds a local full-text and semantic index of it. Questions are then answered from the papers, manuals and documentation, not from memory.
Nothing is invented on top: the server reads the README, the section drafts and the benchmarks at start-up, the README stays the product, and no generated index is committed anywhere.
Connect it
Your AI client starts the server itself, with uv:
it runs uvx cpu-perf, which fetches the release and starts it. Add that
command to your client once, as below; run in a terminal, it only waits for
a client. (pip install cpu-perf works too and gives the same cpu-perf
command.)
The first start downloads its dependencies, which can take longer than some
clients wait for a new server. Run uvx cpu-perf --version once in a
terminal first, and every client after that starts it in about a second.
Every client above starts the server on your own machine; there is nothing to host. Claude.ai, the mobile apps and ChatGPT on the web connect only to servers on the internet, not to a program on your computer, so they cannot use it yet.
Claude
Claude Code
Claude Desktop: Settings, Developer, Edit Config, then add to
claude_desktop_config.json. Desktop does not always see your shell's
PATH, so give the full path that which uvx prints:
ChatGPT
The ChatGPT desktop app runs local servers through its Codex host, configured as below.
Codex
or, in ~/.codex/config.toml (shared by the CLI, the IDE extension and the
ChatGPT desktop app):
Cursor and VS Code
Cursor (~/.cursor/mcp.json) and VS Code (.vscode/mcp.json,
which names the key servers and adds "type": "stdio"):
What every client sees
The answers are markdown written for a model to read. Claude Code and Codex
show the model only a tool's structured data when a tool returns any, so the
server returns none by default and every client reads the same text
(--structured-output adds it back for programmatic use). Every tool is
marked read-only, so no client asks for approval on each call, and slow
reads of large documents return within a minute, finishing in the
background.
Then ask, for example:
- Why does my multithreaded counter stop scaling past two threads?
- Here is my
perf statoutput; where is the time going? - gcc says "not vectorized: complicated access pattern"; what do I change?
- What does
cycle_activity.stalls_l3_misscount? - What does the Intel optimisation manual say about store forwarding?
- Give me a reading path for NUMA, ending with something I can run.
- Why is cppreference not in the list?
- Is "AVX-512 gives 2x on Zen 5" a claim the list would quote?
Use it for your own work
ask takes the question and, optionally, whatever the user pasted as
context. The server reads that output itself, with no model involved:
perf statin its plain,-xand-jforms, per-CPU and interval output included: the counters as read, and the ratios computed from them (instructions per cycle, frequency, branch, cache and TLB miss rates, misses per thousand instructions, stalled-cycle shares, faults and context switches per second), each with its formula, computed as perf computes its own columns. P-core and E-core counts on hybrid parts are never divided by each other.- Top-down level 1 from
perf stat --topdown,-M TopdownL1, the AMDPipelineL1group, or toplev, and level 2 from-M TopdownL2. A level is flagged only against a threshold a listed source states: Intel's own values from its TMA metrics sheet, applied only to Intel P-cores and cited with every flag. Everything else is reported as measured, without a verdict. A flagged level sends the answer to the part of the list about it: a memory-bound run to the memory hierarchy, a front-end-bound one to fetch and decode. - The machine: when the PMUs and event names show Intel, AMD or Arm, sources about the other vendors' hardware and tools are left out.
- How far to trust it: multiplexed counters (and the lowest running
share, metric groups included), events that were not counted or not
supported, how many
-Iintervals were summed, and lines the parser could not read, which are listed rather than guessed. - perf's own errors: a missing metric group,
perf_event_paranoidrefusals, unknown or unsupported events, the NMI watchdog: each restated with what perf itself says to do, instead of being searched for word by word. - gcc
-fopt-infoand clang-Rpassremarks: why each loop was left scalar, counted per loop, routed to the list's auto-vectorisation sources and benchmark. - Assembly and code:
objdump -dwith or without the opcode bytes, gdb'sdisassemble,perf annotate, and source code. The instructions and identifiers that matter (gathers, atomics, fences, intrinsics,alignas,restrict) are routed to the matching sections, and so is what a loop does: a float sum carried across iterations, which stays one serial chain of adds without-fassociative-matheven when a remark says the loop was vectorised (and packed multiplies feeding a run of scalar adds, its shape in assembly), an early exit, or fields read from an array of structs.
The event names, remarks and identifiers then steer the search, so the passages that come back are about the pasted output, not just the question.
Every answer says which of the list's sources it carries text from and which
it does not. Each entry is marked as quoted (with the passages and pages),
in the library but without a matching passage (with the read_source call that
looks inside), or not read on this machine (with the reason: blocked, refused
by a proxy, not fetched yet, a talk with only its description). The model is
told to attribute a claim to a source only through a passage it was given, and
to offer an unread source as further reading, never as a citation. A listed
paper the question is about comes with its abstract, and every passage carries
the list's title for its source rather than the document's own, which is often
a placeholder such as "Untitled Document".
Answers are brief by default: passages are trimmed to the part that matches,
the benchmark and the editorial record come only when they are relevant, and
each passage carries an id that read_source(ref, passage=id) expands in
full. detail="full" returns whole passages and everything related.
Repeated questions are answered from a cache until the library changes.
The first run: building the library
The server answers from the repository immediately. In the background it fetches the linked sources into a local library:
- where:
~/.local/share/cpu-perfon Linux,~/Library/Application Support/cpu-perfon macOS,%LOCALAPPDATA%\cpu-perfon Windows, orCPU_PERF_DATA_DIR; - how long: a few minutes to a quarter of an hour, depending on the network and the large manuals; the crawl resumes where it stopped if the client closes the server;
- how big: one SQLite file of passages, a keyword index and embeddings, a few hundred megabytes at most; the downloaded files themselves are not kept;
- what it skips: very large PDFs are indexed up to a page cap and the rest is read on demand; scanned PDFs, compressed PostScript and videos have no text to index (videos keep their title and description).
To build it in the foreground with progress, run
cpu-perf index; cpu-perf status --detail lists every source with
its state.
Several MCP clients (Claude Desktop, Claude Code and Cursor at once, say) share one library. Every few minutes one of them, whichever holds an operating-system lock on the data folder, does the upkeep: it fetches sources that are new, due for a refresh or due for a retry, embeds passages that have no vector, and brings a library built by an older release up to date in place; a source is fetched again only when a release improves how its kind of document is read (this one rejoins words PDFs hyphenate across lines). A refresh that fails keeps the copy already in the library, unless the document is gone. The lock is released by the system if that client exits or crashes, and the work pauses while requests arrive.
Some publishers (ACM, IEEE, parts of the Intel and Arm portals) refuse automated clients or render their documents only in a browser. Those sources are reported as blocked or partial, with the list's own link notes on why, and the answer points the reader to the link instead. Coverage is reported as it is, never padded.
Semantic search uses the small static embedding model
potion-base-8M, downloaded
once. Without it (offline, or CPU_PERF_EMBED_MODEL=none) the library
falls back to keyword search alone.
Copyright and politeness
The library is built on the user's own machine, or the operator's own server,
from the public URLs the list links; nothing crawled is committed, published
or shipped in the package or the container image. The crawler fetches only
those documents, identifies itself, waits between requests to one host and
backs off on rate limits. It reads each listed link the way a reader opening
it would, so it does not consult robots.txt unless asked to with
--respect-robots (CPU_PERF_RESPECT_ROBOTS=1); sources a site then
disallows are reported as blocked.
Tools
Resources: cpuperf://readme, cpuperf://contents, cpuperf://rules,
cpuperf://watchlist, cpuperf://benchmarks, and the templates
cpuperf://section/{number}, cpuperf://entry/{id},
cpuperf://benchmark/{slug}, cpuperf://file/{+path} and
cpuperf://source/{id}.
Prompts: ask_the_list, study_plan, diagnose (the list's own method:
the USE method, counters that work, top-down, roofline, then the mechanism),
audit_claim, reproduce_benchmark and review_candidate (pre-screens a
proposed entry against CONTRIBUTING.md).
Entry ids (4.3.5, Start here 1.7) follow the README and change when it
does; every output also carries the URL, which does not.
Keeping the list current
Once a day (one request shared by every client on the machine) the server
asks GitHub for the newest commit of the list. When there is a newer one it
downloads that commit, keeps only the list's own files (the same set the
package bundles), checks that they parse into a list no smaller than the one
it is serving, and switches to it between two calls. Links the new list adds
are fetched by the next upkeep pass; links it drops leave the answers. An
entry id that now names a different source is flagged in get_entry.
Downloaded files are read as data. Nothing from them is imported or run, file
sizes and paths are checked before anything is written, and a copy that does
not parse is kept off with a note in library_status to upgrade the server.
A checkout (--repo, or running from the repository) is never updated: it is
the copy being edited. CPU_PERF_AUTO_UPDATE=0 turns the check off;
CPU_PERF_UPSTREAM=owner/repo follows a fork instead.
How it stays honest
- It reads the README with the same grammar as
misc/scripts/check_format.py; a test fails if the two drift apart, and another fails if the parsed counts disagree with the README's badges or the changelog's totals. - The README is authoritative. The section drafts contribute only their Rejected, Claims, Link notes and Benchmark proposal blocks, joined by file number.
- Benchmark numbers come from one Apple M4 Pro; outputs say so, and the server tells the client to quote a number only with all seven fields.
- Text from sources is fenced and labelled as untrusted data.
- Metrics from pasted output are computed exactly as the output shows them, and the only thresholds applied are the ones Intel publishes for top-down level 1, cited each time.
- A retrieval test set of everyday questions guards the ranking: every change
must keep its recall (
python tests/eval_queries.py path/to/library.sqliteprints the report).
Safety of fetching
Only URLs that appear in the repository are read. Every request and every
redirect is checked: http and https on their default ports only, and the host
must resolve to public addresses (no loopback, private, link-local or cloud
metadata addresses); the connection is pinned to the address that was checked.
Downloads and decompression are size-capped and time-boxed. HTTPS_PROXY and
NO_PROXY are honoured.
Speed
Everything about the repository is held in memory and every lookup is a
dictionary walk; passages are served from SQLite FTS5 and a matrix of
quantised embeddings. Run cpu-perf --selftest to see load, index and
query times on your own machine.
Configuration
Serving over HTTP
Nobody using cpu-perf needs this: every client above starts it locally. It is for running one shared instance. The same server speaks Streamable HTTP; from the repository root:
The endpoint is /mcp, with /healthz for health checks, and goes behind
HTTPS. Without an allowed host the server refuses requests addressed to any
host but localhost, which is what protects it from DNS rebinding. There is
no authentication, so anyone with the URL can call the tools, all of which
only read. The container builds its library into the /data volume on first
start; the image itself carries no crawled text. A public instance serves
passages of other people's work alongside their links, much as a search
engine shows snippets.
Development
An editable install reads the checkout, so a README edit shows up on the next
start. The tests need no network: the crawler runs against a local fixture
site and a deterministic embedder, and the daily update against a local
stand-in for GitHub. With CPU_PERF_EVAL_DB pointing at a crawled
library.sqlite, the retrieval tests also run against real sources.
Releasing
Set the version in pyproject.toml, merge, then tag the merge commit on
main with the same version:
.github/workflows/mcp-release.yml builds, tests and publishes to PyPI with
trusted publishing. Once, before the first release: on PyPI add a pending
publisher for project cpu-perf, owner usamahz, repository
cpu-performance-engineering, workflow mcp-release.yml, environment
pypi. GitHub creates the pypi environment on the first run.
A release candidate (0.1.1rc1, say) is tagged the same way; it installs
only when asked for by version, with uvx [email protected], so testing one
never reaches people on the latest release.
The same workflow then lists the release in the
MCP Registry from server.json,
signing in with the workflow's own GitHub identity, so a release needs no
other step. server.json carries the same version as pyproject.toml. The
registry proves the PyPI package belongs to the listing by finding this line
in the README as PyPI shows it, so it stays here:
Licence
MIT, as the repository. The wheel carries the repository's files and its LICENSE.
来源:misc/mcp/README.md,提交 deb5a0b
工具
0版本历史
1- v0.1.1最新Oct 5, 2026

