sota-anchor

io.github.luckmanqasimv0.2.0更新於 Oct 4, 2026

Finds what shipped after your coding agent's training, before it rebuilds it from scratch.

概覽

AI 產生的概覽

讓程式助理在斷定某個函式庫、讀取器或模型不存在之前,先檢索過去一年的論文、程式碼儲存庫與套件。

功能
當計畫建立在「某樣東西不可用」的假設上時,sota-anchor 會執行三步檢查:助理先說出該假設,sota-anchor 檢索最近 12 個月的 arXiv、Hugging Face、GitHub、npm 與 crates.io(設定 Brave 金鑰後還包括網頁),接著助理只根據回傳結果判斷。它提供 check_what_exists 工具與 models://active 資源,工作階段啟動鉤子會加入一份帶日期的目前可用模型 API ID 清單。命令列也提供 check、evidence、sync 與 serve 等指令。
適用情境
當助理因為沒聽過某個函式庫、格式讀取器或模型,而準備從頭重寫或聲稱任務無法完成時,適合使用。它面向 Claude Code、Cursor、VS Code Copilot、Codex CLI、Gemini CLI、OpenCode 等程式助理,也可用於 CI 閘門。
執行需求
本機程序:需要 Python 3.11+ 與 uv(或 pipx),從 PyPI 安裝 sota-anchor,以 uvx 或 sota-anchor serve 執行。一般由助理判斷的檢查不需要 API 金鑰。選用密鑰:SOTA_ANCHOR_BRAVE_API_KEY 用於網頁搜尋,SOTA_ANCHOR_GITHUB_TOKEN 用於提高 GitHub 搜尋額度。CI 無介面判定需要 SOTA_ANCHOR_API_KEY,可選 SOTA_ANCHOR_BASE_URL 與 SOTA_ANCHOR_MODEL。需要連線至公開搜尋 API 的網路。
安裝前請注意
檢查會把助理撰寫的查詢傳送到公開 API(arXiv、Hugging Face、GitHub、npm、crates.io,設定金鑰時還有 Brave);帶權限範圍的 GitHub 權杖可能把私有儲存庫帶進結果,應使用沒有權限範圍的權杖。若設定了 SOTA_ANCHOR_API_KEY,計畫及其證據會傳送到 SOTA_ANCHOR_BASE_URL 指定的 API。sync 指令會在目前專案的 CLAUDE.md、AGENTS.md 或 Cursor 規則中寫入帶標記的區塊。檢索結果是線索而非結論,檢索在 60 秒時停止。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 sota-anchor,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

sota-anchor

Your coding agent doesn't know what shipped after its training.
sota-anchor is an MCP server and Claude Code plugin that checks the last year of papers,
repositories and packages before your agent rebuilds something that already exists,
or tells you it can't be done.

[CI] [License: MIT] [Python 3.11+] [MCP server] [Claude Code plugin] [No API key needed]

[Two workflows from the same task, parsing Xcode 27 project files, to the same goal, build settings checked in CI. Without sota-anchor, the agent rebuilds it from what the model knew: strip trailing commas, model the project, resolve build settings, write a rule engine, write the CLI and tests; the model never saw Apple's library, so that is 755 lines of new code in 2 minutes 37 seconds. With sota-anchor, it checks what shipped since training, finds xcode-project-format, Apple's own library for the format, released on 2026-09-15, after the model's training, and reuses it; no code written, 42 seconds.]

A coding agent plans from what it learned in training. When a task needs something it hasn't heard of, it assumes that thing doesn't exist, and does one of two expensive things: builds it from scratch, or tells you it can't be done. Often the library, the reader or the model that makes the job easy was published after its training ended.

sota-anchor makes the agent check what's been published since. When a plan rests on something being unavailable, the agent searches the last year of papers, repositories and packages, and weighs what comes back, with dates. Then it reuses what exists, or builds knowing that nothing better does.

  • Any field. Nothing in the code knows about any domain. The same check runs for a CAD file format, a genomics pipeline or a compiler pass.
  • No extra API key. Your agent does the reasoning; sota-anchor does the retrieval.
  • Evidence, not memory. Every claim cites a dated source. When the search finds nothing, it says so instead of guessing.

Before and after

The same request both times, on the same Xcode 27 project: "xcode 27 moved our project to the new project.xcproj format and theres no library for it yet, so write a swift parser we can use in CI to check each target's build settings".

Before. Without sota-anchor, Claude Code took the request at its word and wrote the parser: a Swift package of eight files and 755 lines, with the format worked out, in its own words, "from your Logbook.xcodeproj/project.xcproj, not from any Apple documentation". It took 2 minutes 37 seconds.

[A Claude Code session in a Mac terminal. Asked for a Swift parser for Xcode 27's project.xcproj format, which the request says has no library yet, Claude Code writes one from scratch: Package.swift, a trailing-comma stripper, a project model of 147 lines, a build-settings resolver of 169 lines, and more files below the fold.]

After. With sota-anchor, the check runs before any code. It finds Apple's own library for the format, xcode-project-format, published under Apache 2.0 on 2026-09-15, after the model was trained. Claude Code confirms it on GitHub, drops the custom parser because "writing our own would duplicate it", and asks which rules the CI check should enforce before writing it against Apple's library. 42 seconds, and no code written yet.

[The same request with sota-anchor. Claude Code calls the sota-anchor check twice and fetches github.com/apple/xcode-project-format. It reports that Apple publishes an official Swift library for the format, says writing its own parser would duplicate it, and asks which rules the CI check should enforce. Done in 42 seconds, with no code written.]

Both are recorded Claude Code sessions (Opus 5.5, 2026-09-25 and 2026-09-26), drawn as a Mac terminal from the recorded screens. The first is shown from the top and runs on for another 116 rows. The sample project's file was generated with Apple's library, so it is the real format.

The same check on other plans

Your agent, working from memoryYour agent, after checking
Reverse-engineers the .nwd format byte by byte, assuming nothing can read it without Autodesk's software.Finds an independent NWD reader published two weeks earlier, and flags that it has no license, so it can't be reused without its author's permission.
Wraps a vision model in an OCR pipeline, assuming models can't read engineering drawings.Brings back this year's benchmarks of multimodal models on exactly those drawings, AECV-Bench and Enginuity, so the choice rests on measured results.
Writes gemini-3-pro into a new project's .env. That was never a served model ID.Uses gemini-3.1-pro-preview, which a public model registry lists as served today.

The first and last rows are from recorded runs. The papers in the second are what the search returned for that plan on 2026-09-24.

How it works

[Your request goes to your agent, which inverts it into what the plan assumes doesn't exist and writes two queries. sota-anchor searches the last year of arXiv, Hugging Face, GitHub, npm and crates.io and returns dated evidence. Your agent judges the evidence, never its memory, and reaches one of three outcomes: reuse what exists, build it, or no verdict.]

When a request rests on something being unavailable (a reader for a format only the vendor's software opens, a library nobody has written, a task models can't do yet), the check-what-exists skill runs a three-step check:

  1. Invert. Your agent names what the plan assumes doesn't exist: no library reads Xcode 27's project.xcproj format yet. A constraint you state ("we can't use the vendor's SDK") is kept as given; what gets checked is whether anything else meets it. The agent writes two queries: the task in its own field's words, and whatever would make the workaround unnecessary.
  2. Search. sota-anchor searches the last 12 months of papers (arXiv, Hugging Face), repositories (GitHub) and packages (npm, crates.io), plus the web if you give it a Brave key. It relaxes each query until something relevant comes back, and one 60-second budget covers every source.
  3. Judge. Your agent weighs only what came back. The assumption stands unless the evidence documents otherwise. "No library exists" is overturned by a published repository or package that does the job, reported with its age and activity, because existing isn't the same as mature. "Models can't do this" needs benchmarked results.

The verdict is one of three. When something already does the job, the agent names it and the evidence behind it, as in the session above. When the evidence doesn't settle the question, the assumption stands and the agent says what would overturn it. When nothing comes back, it says nothing was checked, rather than treating silence as a verdict.

Also: current model IDs

A smaller convenience, mostly for new projects. Every Claude Code session starts with a short, dated list of the model API IDs that OpenRouter's public registry serves today, and the older IDs they replaced, so a fresh config names a model that exists:

diff
- GEMINI_MODEL=gemini-3-pro            # recalled from training: not a served ID+ GEMINI_MODEL=gemini-3.1-pro-preview  # from the session's registry snapshot

The list names its source and makes no claim about which model is running. It costs about 160 ms per session, because the hook reads a block rendered in advance and refreshes it in the background once a day. Other agents get the same list through sota-anchor sync.

Quick start: Claude Code

[!NOTE] Needs Python 3.11+ and uv. No API key.

In Claude Code:

text
/plugin marketplace add luckmanqasim/sota-anchor/plugin install sota-anchor@sota-anchor

Then start a new session. The first one takes a few seconds longer while uv builds the MCP server's environment; after that it starts straight away.

To try it for one session without installing:

bash
git clone https://github.com/luckmanqasim/sota-anchorclaude --plugin-dir ./sota-anchor

What you get:

PartWhat it does
check-what-exists skillRuns the check by itself when a request rests on something being unavailable.
/sota-check <plan>Runs the check on demand.
MCP serverThe check_what_exists tool and the models://active resource.
Session-start hookAdds the dated model list to every session. Nothing to remember to run.
/sota-syncRefreshes the model list now and writes it into CLAUDE.md, AGENTS.md and Cursor's rules.
Prompt hook (opt-in)With SOTA_ANCHOR_PROMPT_HOOK=1, adds a nudge when a message says something can't be done.

Other coding agents

The check is a plain MCP server, so it works in any agent that speaks MCP. The model list goes into the project's instructions file.

AgentThe checkModel list comes from
Claude Codeautomatic, or /sota-checkthe plugin, every session
Cursorcheck_what_exists tool.cursor/rules/sota.mdc
VS Code + GitHub Copilotcheck_what_exists toolAGENTS.md
OpenAI Codex CLIcheck_what_exists toolAGENTS.md
Gemini CLIcheck_what_exists toolAGENTS.md
OpenCodecheck_what_exists toolAGENTS.md
Any other MCP clientcheck_what_exists toolyour client's instructions file

1. Install the CLI once:

bash
uv tool install sota-anchor   # or: pipx install sota-anchor

2. Register the MCP server in your agent, using the snippet for it below.

3. Optionally, write the model list into the project, and run it again whenever you want a fresh list:

bash
sota-anchor sync --target agents   # writes AGENTS.mdsota-anchor sync --target cursor   # writes .cursor/rules/sota.mdc and .cursorrules

Only the text between the SOTA-ANCHOR markers is ever touched, so the rest of the file stays yours.

[!TIP] Outside Claude Code, your agent decides when to call check_what_exists from the tool's own description. To be sure it runs, ask for it: "run check_what_exists on this plan before building it."

Cursor

.cursor/mcp.json in the project, or ~/.cursor/mcp.json for every project:

json
{  "mcpServers": {    "sota-anchor": { "command": "sota-anchor", "args": ["serve"] }  }}

Model list: sota-anchor sync --target cursor writes .cursor/rules/sota.mdc, which is always applied. Cursor also reads AGENTS.md.

VS Code (GitHub Copilot)

.vscode/mcp.json. VS Code's key is servers, not mcpServers:

json
{  "servers": {    "sota-anchor": { "command": "sota-anchor", "args": ["serve"] }  }}

Or add it to your user profile from a terminal:

bash
code --add-mcp '{"name":"sota-anchor","command":"sota-anchor","args":["serve"]}'

Model list: sota-anchor sync --target agents, then turn on the chat.useAgentsMdFile setting. VS Code's local agent doesn't read AGENTS.md by default.

OpenAI Codex CLI
bash
codex mcp add sota-anchor -- sota-anchor serve

or in ~/.codex/config.toml:

toml
[mcp_servers.sota-anchor]command = "sota-anchor"args = ["serve"]

Model list: sota-anchor sync --target agents. Codex reads AGENTS.md.

Gemini CLI
bash
gemini mcp add sota-anchor sota-anchor serve

or in ~/.gemini/settings.json (.gemini/settings.json for one project). That file is also where you tell Gemini CLI to read AGENTS.md, since it reads GEMINI.md by default:

json
{  "mcpServers": {    "sota-anchor": { "command": "sota-anchor", "args": ["serve"] }  },  "context": { "fileName": ["AGENTS.md", "GEMINI.md"] }}

Model list: sota-anchor sync --target agents.

OpenCode

opencode.json in the project:

json
{  "$schema": "https://opencode.ai/config.json",  "mcp": {    "sota-anchor": { "type": "local", "command": ["sota-anchor", "serve"], "enabled": true }  }}

To run it through uvx instead of installing it, raise the timeout: OpenCode waits 5 seconds for a server's tools, and the first launch spends longer than that downloading packages.

json
"sota-anchor": {  "type": "local",  "command": ["uvx", "sota-anchor", "serve"],  "enabled": true,  "timeout": 30000}

Model list: sota-anchor sync --target agents. OpenCode reads AGENTS.md, and falls back to CLAUDE.md when there isn't one.

Any other MCP client

Most clients, Claude Desktop included, take the common mcpServers shape. Check your client's docs for where the file lives.

json
{  "mcpServers": {    "sota-anchor": { "command": "sota-anchor", "args": ["serve"] }  }}

To run it without installing anything, let uv fetch it from PyPI on demand. The first launch downloads its packages, so a client with a short startup timeout may need a second try:

json
{  "mcpServers": {    "sota-anchor": { "command": "uvx", "args": ["sota-anchor", "serve"] }  }}

In Claude Code without the plugin: claude mcp add sota-anchor -- sota-anchor serve.

[!TIP] If your editor says it can't find sota-anchor, give it the full path. uv tool dir --bin prints the folder uv installed it into.

Command line

Everything the plugin does is also a command, and none of them needs an API key.

CommandWhat it does
sota-anchor check "<your plan>"Run the whole check. See In CI for what it returns.
sota-anchor evidence --query "Xcode 27 project.xcproj parser"Search for evidence. No LLM involved; --json for scripts.
sota-anchor serveRun the MCP server over stdio.
sota-anchor seedPrint the session block. --refresh fetches the registry first.
sota-anchor sync --target allWrite the model list into CLAUDE.md, AGENTS.md and Cursor's rules.

sync options:

OptionWhat it does
--targetclaude (the default), cursor, agents or all.
--providerWhich providers to list. Repeatable; defaults to anthropic, google and openai.
--all-providersList every provider: about 4,600 tokens of context against 510.
--max-per-providerEndpoints listed per provider. Default 4.
--staleness-monthsHow far behind its provider's newest release a model can fall before it counts as superseded. Default 12; raise it for providers that ship slowly.
--refreshIgnore the 24-hour cache.

In CI

Without an API key, check prints the protocol for an agent to answer. With one it reaches a verdict itself, which makes it usable as a gate:

Exit codeMeaning
0The limitation still holds.
2The plan relies on something obsolete.
3No verdict: no key was set, so an agent still has to judge. Deliberately not 0.
yaml
- run: uv tool install sota-anchor- run: sota-anchor check "$(cat docs/design-notes.md)"  env:    SOTA_ANCHOR_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}

Configuration

Nothing needs to be set. These are all optional.

In Claude Code, the plugin asks for its two optional keys when you enable it, and keeps them in your system's credential store. Leave either empty to go without.

Plugin settingEffect
GitHub tokenRaises GitHub's search limit from 10 to 30 requests a minute. Give it no scopes: a token that can see your private repositories adds them to the search results your agent reads.
Brave Search API keyAdds general web search as an evidence source. Unset, the web is not queried.

sota-anchor only reads variables named for it, so a key you've set for other tools, such as GITHUB_TOKEN, is never picked up. Outside the plugin, set these instead:

VariableEffect
SOTA_ANCHOR_GITHUB_TOKENThe GitHub token above.
SOTA_ANCHOR_BRAVE_API_KEYThe Brave Search key above.
SOTA_ANCHOR_CACHE_DIRWhere the catalog and session block are cached. Default ~/.cache/sota-anchor.
SOTA_ANCHOR_TTL_MINUTESHow old the session block can get before the hook refreshes it. Default 1440.
SOTA_ANCHOR_PROMPT_HOOK1 turns on the prompt-time nudge in Claude Code.

For a headless verdict, where no agent is present to judge (as in CI):

VariableEffect
SOTA_ANCHOR_API_KEYAny OpenAI-compatible key.
SOTA_ANCHOR_BASE_URLThe API base URL. Defaults to OpenRouter; set it to use another provider's key. Must be https://, or http:// for a server on this machine (localhost, 127.0.0.1, ::1).
SOTA_ANCHOR_MODELThe judging model. Unset, it is picked from the live catalog.

What it sends and fetches

Everything sota-anchor does over the network, and every file it writes. There is no telemetry.

When the check runs, the search queries your agent writes from your request, and shorter versions of them as the search relaxes, go to these public APIs with a sota-anchor/<version> User-Agent:

ServiceEndpointWhat it receives
arXivexport.arxiv.org/api/querythe queries
Hugging Facehuggingface.co/api/papers/searchthe queries
GitHubapi.github.com/search/repositoriesthe queries, and your GitHub token if you set one
npmregistry.npmjs.org/-/v1/searchthe queries
crates.iocrates.io/api/v1/cratesthe queries
Brave Searchapi.search.brave.com/res/v1/web/searchthe queries and your key, only if you set one

For the model list, it fetches openrouter.ai/api/v1/models, a public endpoint that takes no key and receives nothing about you. The session hook refreshes it in the background at most once a day.

For a headless verdict, when SOTA_ANCHOR_API_KEY is set, the plan and the evidence found for it go to the API at SOTA_ANCHOR_BASE_URL, which is openrouter.ai/api/v1 unless you change it. Both sota-anchor check and the MCP tool do this, including the plugin's server if the variable is in the environment Claude Code starts from. Without it, your agent does the judging and nothing goes to an LLM API.

On first launch, uv downloads the Python packages the server needs from PyPI, at the exact versions uv.lock pins.

On disk, it writes the model catalog and the session block to ~/.cache/sota-anchor, or to SOTA_ANCHOR_CACHE_DIR. /sota-sync and sota-anchor sync write a marked block into CLAUDE.md, AGENTS.md or Cursor's rules in the current project, and only when you run them. The opt-in prompt hook reads your prompt on your machine to decide whether to add a nudge, and sends it nowhere.

Limitations

  • Whether the check runs is your agent's decision. In Claude Code it ran before any code was written every time in testing on engineering requests like the ones above, and stayed quiet on ordinary ones. In other agents, ask for it by name when it matters.
  • A find is a lead, not a decision. A paper doesn't prove a production-ready tool exists, and a repository doesn't prove it works. Treat a find as a reason to look.
  • Relevance is lexical. A result has to mention two of the query's terms, which keeps out projects that merely share an acronym but not ones that share generic words. Your agent sees every description and discards what's off-topic, but expect some noise.
  • PyPI isn't searched. It has no search API. Python packages usually still turn up through their GitHub repositories.
  • arXiv is slow and particular. Requests are spaced 3.5 s apart per its terms of use. Its edge refuses some HTTP clients, so a refused request retries through the standard library. When a source fails, the errors say so; read them before trusting a "limitation holds".
  • Retrieval stops at 60 seconds. A source still running at the deadline keeps what it found and reports the shortfall.
  • The model list can be a day old. The hook never waits on the network. Run /sota-sync to refresh it immediately.
  • The headless verdict is untested against a live provider. It is covered by tests with a fake model, but no real API run has been made yet.

Design notes

  • No topic dictionaries. Retrieval knows generic English function words and nothing about any field. Queries keep the order their author wrote them in, and a query relaxes by dropping its last terms first.
  • AND, not OR. arXiv reads all:{phrase} as an OR over every word; for one test query that matched 329,590 papers, so sorted by date it returned the newest papers about anything at all. Terms are ANDed, and the query is relaxed only when it finds nothing.
  • No verdict from nothing. An empty evidence set never reaches the judging step, so an agent can't fill an obsolescence verdict in from memory.
  • Reference data, not orders. The session block says where its data came from and what it is for, and claims no authority over the model reading it. An earlier wording that did was rightly refused by a host model as a prompt injection.
  • No hardcoded models. Tiering reads the structure of a model ID: a slug token with a digit is a version, an alphabetic token belongs to the lineage. So gpt-4o and gpt-5.5 share a lineage across a naming change. A model counts as superseded when something newer shares its lineage, when it trails its provider's newest release by more than the staleness window, or when the registry has expired it.

Development

bash
git clone https://github.com/luckmanqasim/sota-anchorcd sota-anchoruv sync --extra devuv run pytest            # all offlineuv run ruff check .claude plugin validate .

The suite never touches the network or your own cache. HTTP goes through httpx.MockTransport and the LLM through an injected fake. Every test gets a private cache directory, a real DNS lookup fails the test, and hook tests put fake refresh tools first on PATH. CI runs it on Ubuntu and Windows with Python 3.11, 3.12, 3.13 and 3.14.

Background

SciUnlearn (Paul, Patwardhan & Cohan, arXiv:2608.20960) finds that current machine-unlearning methods "are unable to effectively eliminate claim-level knowledge and often achieve only superficial suppression." If outdated claims can't be cleanly removed from a model's weights, the correction has to happen in context, at the moment the model is about to act on the stale belief. That is where sota-anchor works.

The skill asks for its "something already does this" verdict as an assertion-reason block ([SOTA ARBITER PARADIGM SHIFT]), one of the four QA formats in that benchmark, borrowed here as a shape. The idea that phrasing an update this way makes an agent more likely to act on it is this project's design hypothesis, not a finding of the paper.

License

MIT © 2026 Luckman Qasim

來源:README.md,提交 6fccdb4

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.2.0最新Oct 4, 2026