
sota-anchor
io.github.luckmanqasimv0.2.0更新於 Oct 4, 2026
Finds what shipped after your coding agent's training, before it rebuilds it from scratch.
概覽
讓程式助理在斷定某個函式庫、讀取器或模型不存在之前,先檢索過去一年的論文、程式碼儲存庫與套件。
- 功能
- 當計畫建立在「某樣東西不可用」的假設上時,sota-anchor 會執行三步檢查:助理先說出該假設,sota-anchor 檢索最近 12 個月的 arXiv、Hugging Face、GitHub、npm 與 crates.io(設定 Brave 金鑰後還包括網頁),接著助理只根據回傳結果判斷。它提供 check_what_exists 工具與 models://active 資源,工作階段啟動鉤子會加入一份帶日期的目前可用模型 API ID 清單。命令列也提供 check、evidence、sync 與 serve 等指令。
- 適用情境
- 當助理因為沒聽過某個函式庫、格式讀取器或模型,而準備從頭重寫或聲稱任務無法完成時,適合使用。它面向 Claude Code、Cursor、VS Code Copilot、Codex CLI、Gemini CLI、OpenCode 等程式助理,也可用於 CI 閘門。
- 執行需求
- 本機程序:需要 Python 3.11+ 與 uv(或 pipx),從 PyPI 安裝 sota-anchor,以 uvx 或 sota-anchor serve 執行。一般由助理判斷的檢查不需要 API 金鑰。選用密鑰:SOTA_ANCHOR_BRAVE_API_KEY 用於網頁搜尋,SOTA_ANCHOR_GITHUB_TOKEN 用於提高 GitHub 搜尋額度。CI 無介面判定需要 SOTA_ANCHOR_API_KEY,可選 SOTA_ANCHOR_BASE_URL 與 SOTA_ANCHOR_MODEL。需要連線至公開搜尋 API 的網路。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 sota-anchor,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
sota-anchor
Your coding agent doesn't know what shipped after its training.
sota-anchor is an MCP server and Claude Code plugin that checks the last year of papers,
repositories and packages before your agent rebuilds something that already exists,
or tells you it can't be done.
[CI] [License: MIT] [Python 3.11+] [MCP server] [Claude Code plugin] [No API key needed]
A coding agent plans from what it learned in training. When a task needs something it hasn't heard of, it assumes that thing doesn't exist, and does one of two expensive things: builds it from scratch, or tells you it can't be done. Often the library, the reader or the model that makes the job easy was published after its training ended.
sota-anchor makes the agent check what's been published since. When a plan rests on something being unavailable, the agent searches the last year of papers, repositories and packages, and weighs what comes back, with dates. Then it reuses what exists, or builds knowing that nothing better does.
- Any field. Nothing in the code knows about any domain. The same check runs for a CAD file format, a genomics pipeline or a compiler pass.
- No extra API key. Your agent does the reasoning; sota-anchor does the retrieval.
- Evidence, not memory. Every claim cites a dated source. When the search finds nothing, it says so instead of guessing.
Before and after
The same request both times, on the same Xcode 27 project: "xcode 27 moved our project to the new project.xcproj format and theres no library for it yet, so write a swift parser we can use in CI to check each target's build settings".
Before. Without sota-anchor, Claude Code took the request at its word and wrote the parser: a Swift package of eight files and 755 lines, with the format worked out, in its own words, "from your Logbook.xcodeproj/project.xcproj, not from any Apple documentation". It took 2 minutes 37 seconds.
After. With sota-anchor, the check runs before any code. It finds Apple's own library for the format, xcode-project-format, published under Apache 2.0 on 2026-09-15, after the model was trained. Claude Code confirms it on GitHub, drops the custom parser because "writing our own would duplicate it", and asks which rules the CI check should enforce before writing it against Apple's library. 42 seconds, and no code written yet.
Both are recorded Claude Code sessions (Opus 5.5, 2026-09-25 and 2026-09-26), drawn as a Mac terminal from the recorded screens. The first is shown from the top and runs on for another 116 rows. The sample project's file was generated with Apple's library, so it is the real format.
The same check on other plans
The first and last rows are from recorded runs. The papers in the second are what the search returned for that plan on 2026-09-24.
How it works
When a request rests on something being unavailable (a reader for a format only the vendor's
software opens, a library nobody has written, a task models can't do yet), the
check-what-exists skill runs a three-step check:
- Invert. Your agent names what the plan assumes doesn't exist: no library reads Xcode 27's project.xcproj format yet. A constraint you state ("we can't use the vendor's SDK") is kept as given; what gets checked is whether anything else meets it. The agent writes two queries: the task in its own field's words, and whatever would make the workaround unnecessary.
- Search. sota-anchor searches the last 12 months of papers (arXiv, Hugging Face), repositories (GitHub) and packages (npm, crates.io), plus the web if you give it a Brave key. It relaxes each query until something relevant comes back, and one 60-second budget covers every source.
- Judge. Your agent weighs only what came back. The assumption stands unless the evidence documents otherwise. "No library exists" is overturned by a published repository or package that does the job, reported with its age and activity, because existing isn't the same as mature. "Models can't do this" needs benchmarked results.
The verdict is one of three. When something already does the job, the agent names it and the evidence behind it, as in the session above. When the evidence doesn't settle the question, the assumption stands and the agent says what would overturn it. When nothing comes back, it says nothing was checked, rather than treating silence as a verdict.
Also: current model IDs
A smaller convenience, mostly for new projects. Every Claude Code session starts with a short, dated list of the model API IDs that OpenRouter's public registry serves today, and the older IDs they replaced, so a fresh config names a model that exists:
The list names its source and makes no claim about which model is running. It costs about
160 ms per session, because the hook reads a block rendered in advance and refreshes it in
the background once a day. Other agents get the same list through sota-anchor sync.
Quick start: Claude Code
[!NOTE] Needs Python 3.11+ and uv. No API key.
In Claude Code:
Then start a new session. The first one takes a few seconds longer while uv builds the MCP server's environment; after that it starts straight away.
To try it for one session without installing:
What you get:
Other coding agents
The check is a plain MCP server, so it works in any agent that speaks MCP. The model list goes into the project's instructions file.
1. Install the CLI once:
2. Register the MCP server in your agent, using the snippet for it below.
3. Optionally, write the model list into the project, and run it again whenever you want a fresh list:
Only the text between the SOTA-ANCHOR markers is ever touched, so the rest of the file
stays yours.
[!TIP] Outside Claude Code, your agent decides when to call
check_what_existsfrom the tool's own description. To be sure it runs, ask for it: "run check_what_exists on this plan before building it."
Cursor
.cursor/mcp.json in the project, or ~/.cursor/mcp.json for every project:
Model list: sota-anchor sync --target cursor writes .cursor/rules/sota.mdc, which is
always applied. Cursor also reads AGENTS.md.
VS Code (GitHub Copilot)
.vscode/mcp.json. VS Code's key is servers, not mcpServers:
Or add it to your user profile from a terminal:
Model list: sota-anchor sync --target agents, then turn on the chat.useAgentsMdFile
setting. VS Code's local agent doesn't read AGENTS.md by default.
OpenAI Codex CLI
or in ~/.codex/config.toml:
Model list: sota-anchor sync --target agents. Codex reads AGENTS.md.
Gemini CLI
or in ~/.gemini/settings.json (.gemini/settings.json for one project). That file is
also where you tell Gemini CLI to read AGENTS.md, since it reads GEMINI.md by default:
Model list: sota-anchor sync --target agents.
OpenCode
opencode.json in the project:
To run it through uvx instead of installing it, raise the timeout: OpenCode waits 5 seconds
for a server's tools, and the first launch spends longer than that downloading packages.
Model list: sota-anchor sync --target agents. OpenCode reads AGENTS.md, and falls back
to CLAUDE.md when there isn't one.
Any other MCP client
Most clients, Claude Desktop included, take the common mcpServers shape. Check your
client's docs for where the file lives.
To run it without installing anything, let uv fetch it from PyPI on demand. The first launch downloads its packages, so a client with a short startup timeout may need a second try:
In Claude Code without the plugin: claude mcp add sota-anchor -- sota-anchor serve.
[!TIP] If your editor says it can't find
sota-anchor, give it the full path.uv tool dir --binprints the folder uv installed it into.
Command line
Everything the plugin does is also a command, and none of them needs an API key.
sync options:
In CI
Without an API key, check prints the protocol for an agent to answer. With one it reaches a
verdict itself, which makes it usable as a gate:
Configuration
Nothing needs to be set. These are all optional.
In Claude Code, the plugin asks for its two optional keys when you enable it, and keeps them in your system's credential store. Leave either empty to go without.
sota-anchor only reads variables named for it, so a key you've set for other tools, such
as GITHUB_TOKEN, is never picked up. Outside the plugin, set these instead:
For a headless verdict, where no agent is present to judge (as in CI):
What it sends and fetches
Everything sota-anchor does over the network, and every file it writes. There is no telemetry.
When the check runs, the search queries your agent writes from your request, and
shorter versions of them as the search relaxes, go to these public APIs with a
sota-anchor/<version> User-Agent:
For the model list, it fetches openrouter.ai/api/v1/models, a public endpoint that
takes no key and receives nothing about you. The session hook refreshes it in the
background at most once a day.
For a headless verdict, when SOTA_ANCHOR_API_KEY is set, the plan and the evidence
found for it go to the API at SOTA_ANCHOR_BASE_URL, which is openrouter.ai/api/v1
unless you change it. Both sota-anchor check and the MCP tool do this, including the
plugin's server if the variable is in the environment Claude Code starts from. Without
it, your agent does the judging and nothing goes to an LLM API.
On first launch, uv downloads the Python packages the server needs from PyPI, at the
exact versions uv.lock pins.
On disk, it writes the model catalog and the session block to ~/.cache/sota-anchor,
or to SOTA_ANCHOR_CACHE_DIR. /sota-sync and sota-anchor sync write a marked block
into CLAUDE.md, AGENTS.md or Cursor's rules in the current project, and only when you
run them. The opt-in prompt hook reads your prompt on your machine to decide whether to
add a nudge, and sends it nowhere.
Limitations
- Whether the check runs is your agent's decision. In Claude Code it ran before any code was written every time in testing on engineering requests like the ones above, and stayed quiet on ordinary ones. In other agents, ask for it by name when it matters.
- A find is a lead, not a decision. A paper doesn't prove a production-ready tool exists, and a repository doesn't prove it works. Treat a find as a reason to look.
- Relevance is lexical. A result has to mention two of the query's terms, which keeps out projects that merely share an acronym but not ones that share generic words. Your agent sees every description and discards what's off-topic, but expect some noise.
- PyPI isn't searched. It has no search API. Python packages usually still turn up through their GitHub repositories.
- arXiv is slow and particular. Requests are spaced 3.5 s apart per its terms of use. Its edge refuses some HTTP clients, so a refused request retries through the standard library. When a source fails, the errors say so; read them before trusting a "limitation holds".
- Retrieval stops at 60 seconds. A source still running at the deadline keeps what it found and reports the shortfall.
- The model list can be a day old. The hook never waits on the network. Run
/sota-syncto refresh it immediately. - The headless verdict is untested against a live provider. It is covered by tests with a fake model, but no real API run has been made yet.
Design notes
- No topic dictionaries. Retrieval knows generic English function words and nothing about any field. Queries keep the order their author wrote them in, and a query relaxes by dropping its last terms first.
- AND, not OR. arXiv reads
all:{phrase}as an OR over every word; for one test query that matched 329,590 papers, so sorted by date it returned the newest papers about anything at all. Terms are ANDed, and the query is relaxed only when it finds nothing. - No verdict from nothing. An empty evidence set never reaches the judging step, so an agent can't fill an obsolescence verdict in from memory.
- Reference data, not orders. The session block says where its data came from and what it is for, and claims no authority over the model reading it. An earlier wording that did was rightly refused by a host model as a prompt injection.
- No hardcoded models. Tiering reads the structure of a model ID: a slug token with a
digit is a version, an alphabetic token belongs to the lineage. So
gpt-4oandgpt-5.5share a lineage across a naming change. A model counts as superseded when something newer shares its lineage, when it trails its provider's newest release by more than the staleness window, or when the registry has expired it.
Development
The suite never touches the network or your own cache. HTTP goes through
httpx.MockTransport and the LLM through an injected fake. Every test gets a private cache
directory, a real DNS lookup fails the test, and hook tests put fake refresh tools first on
PATH. CI runs it on Ubuntu and Windows with Python 3.11, 3.12, 3.13 and 3.14.
Background
SciUnlearn (Paul, Patwardhan & Cohan, arXiv:2608.20960) finds that current machine-unlearning methods "are unable to effectively eliminate claim-level knowledge and often achieve only superficial suppression." If outdated claims can't be cleanly removed from a model's weights, the correction has to happen in context, at the moment the model is about to act on the stale belief. That is where sota-anchor works.
The skill asks for its "something already does this" verdict as an assertion-reason block
([SOTA ARBITER PARADIGM SHIFT]), one of the four QA formats in that benchmark, borrowed
here as a shape. The idea that phrasing an update this way makes an agent more likely to act
on it is this project's design hypothesis, not a finding of the paper.
License
MIT © 2026 Luckman Qasim
來源:README.md,提交 6fccdb4
工具
0版本歷史
1- v0.2.0最新Oct 4, 2026


