Asta MCP: Academic Paper Search
Asta is Ai2's Scientific Corpus Tool: the Semantic Scholar academic graph exposed over MCP (streamable HTTP). This skill tells agents which Asta tool to call for which intent, how to keep responses small enough to read, and which failure modes are silent.
- MCP endpoint:
https://asta-tools.allen.ai/mcp/v1 - Auth:
x-api-keyheader (request key at https://share.hsforms.com/1L4hUh20oT3mu8iXJQMV77w3ioxm) - Transport: streamable HTTP, stateless (no session id, no server-side state between calls)
- Behavior verified against: Asta MCP server v1.12.3, 2026-10-01
Prerequisite Check
Before invoking any tool, verify the Asta MCP server is registered in the host agent. Tool names are prefixed by the MCP server name chosen at install time (commonly asta__<tool> or mcp__asta__<tool>).
If no Asta tools are visible, do not make raw HTTP calls or invent results. Tell the user to register https://asta-tools.allen.ai/mcp/v1 as a streamable HTTP MCP server with an x-api-key header, then restart/reload the host. Minimal setup hints:
- Codex CLI: add
[mcp_servers.asta] url = "https://asta-tools.allen.ai/mcp/v1"andenv_http_headers = { "x-api-key" = "ASTA_API_KEY" }to~/.codex/config.toml. - Claude Code: run
claude mcp add -t http -s user asta https://asta-tools.allen.ai/mcp/v1 -H "x-api-key: $ASTA_API_KEY". - Generic MCP clients: configure server URL
https://asta-tools.allen.ai/mcp/v1with header{ "x-api-key": "<YOUR_API_KEY>" }.
Tool Map: Intent → Asta Tool
Parameter names differ per tool (and unknown parameters are ignored, not rejected)
The server accepts unknown arguments silently, so a wrong parameter name produces an unfiltered but successful call. If results look unfiltered, suspect the parameter name before doubting the corpus.
get_author_papersselects fields withpaper_fields, notfields. Passingfields=is ignored and you get titles only.get_citationshas novenues. Passingvenues=is ignored and citer rows come back unfiltered.snippet_searchhas neitherfieldsnorpublication_date_range. It hasinserted_before(corpus insertion date, formatsYYYY-MM-DD/YYYY-MM/YYYY),paper_ids(comma-separated, up to 100 ids, to restrict snippets to specific papers), andvenues.- Every tool with a
fields/paper_fieldsparameter defaults totitle; omitting it is the quiet way to get an unrankable, unreadable result. - No tool returns a total count of matches. Report the returned row count, never a corpus total.
Row and payload budget (measured sizes)
- Always pass an explicit
limit: 20–50 for paper and citation lists, 5–10 for snippets. Defaults of 50, 100 and 1000 are context blowups. snippet_searchis the one tool whose documented default disagrees with the live server: the published docs page says 250, the live schema says 20, and an unsetlimitmeasured 20 rows / ~85 KB. Trust the live schema and setlimitexplicitly.- Never request
citationsorreferencesthroughfieldsonget_paper/get_paper_batch: one highly cited paper returns 200 KB+. Useget_citationsfor forward citations (it paginates); for a reference list, requestreferencesonly on papers you know have fewer than ~100. - Never request
papersonsearch_authors_by_name: it returns the author's entire publication list and ignoreslimit. Useget_author_paperswith an explicitlimitinstead. - Drop
abstractwhen the output is a table or a ranking pass; fetch it only for the few papers you will actually describe.
Use task-specific field presets.
Metadata lookup:
Search/ranking/results tables:
DOI/export handoff:
Add journal, publicationDate, fieldsOfStudy, isOpenAccess only when needed.
Undocumented fields that do work (pass-through)
externalIds, citationCount and influentialCitationCount are absent from the documented field list but are transparently passed through to Semantic Scholar and return correctly (verified on v1.12.3). externalIds keys seen: DOI, ArXiv, PubMed, PubMedCentral, CorpusId, MAG, DBLP. Caveats:
- Not every paper has a DOI, especially arXiv preprints, which may carry only
ArXiv+CorpusId. - DOI lookup works for the ids tested (
DOI:10.1038/s41586-021-03819-2,DOI:10.1101/...,DOI:10.1371/journal.pone.0000308all resolved), but treat it as best effort and fall back to a title search plusexternalIds. get_paperalso returnspaperId,isOpenAccessandopenAccessPdfwhen not requested;journalis often only{"pages": ...}.- Treat all of this as best effort and degrade gracefully if a future Asta release drops it.
Date and venue filters
publication_date_range accepts YYYY-MM-DD:YYYY-MM-DD with both terms optional, and year or month shorthand: 2021:, :2015-01, 2015:2020, 2019-03. It works on search_papers_by_relevance, search_paper_by_title, get_citations and get_author_papers (verified correct for get_author_papers: 2024:2024 returned exactly the 7 in-range papers of an 87-paper author).
- Papers with an unknown exact date still keep
year, whilepublicationDatecomes backnull. Unknown dates are treated as January 1 of their year, so boundary-year filtering is approximate. venuesmatches Semantic Scholar venue strings. Common abbreviations resolve (verified:NeurIPSandICMLboth returned hits), and thevenuefield comes back canonical (Neural Information Processing Systems). Safest pattern: run one unfiltered keyword search, read thevenuevalues, then filter with those exact strings.- A venue string that matches nothing returns zero rows, not an error, so an empty result can mean a wrong venue string rather than a missing paper.
Response Shapes (verified)
authorsis a list of objects in paper tools but a list of plain strings inside snippets.- Snippet papers carry
corpusId, notpaperId; look one up withCorpusId:<corpusId>. search_paper_by_titleaddsmatchScore; higher is better, useful for accepting or rejecting a fuzzy title.
Workflow Patterns
Pattern 1: Topic Discovery
search_papers_by_relevance(keyword, publication_date_range="<current_year-5>:", fields="title,year,authors,venue,tldr,url,citationCount,influentialCitationCount", limit=20), computing the lower bound from today's date (in 2026, pass2021:); drop or widen the filter if the user wants older work- Rank and present the top N by
citationCount+ recency - Offer follow-ups:
get_citationson the most influential hit, orsnippet_searchfor specific claims
Pattern 2: Seed-Paper Expansion
get_paper(DOI|ARXIV|...)to verify the seedget_citations(paperId, fields="title,year,venue,citationCount", limit=50)for forward expansion- Optionally
search_papers_by_relevancewith the seed's title terms for sideways discovery - Deduplicate by
paperIdbefore presenting
Pattern 3: Author Deep-Dive
search_authors_by_name(name, fields="name,affiliations,paperCount,citationCount,hIndex,externalIds", limit=10). These fields must be requested: the default returns onlyname+authorId. Disambiguate byexternalIds.ORCIDwhen present, thenpaperCount/citationCount/hIndex;affiliationsis frequently empty even when requested (verified) and is only a tiebreaker. Name variants split the same person across ids (F. TheisandFabian J. Theiswere separate profiles), so check the top few candidates rather than the first row.get_author_papers(authorId, limit=50, paper_fields="title,year,authors,venue,tldr,url,citationCount", publication_date_range=?)for a bounded first page, newest first- Filter client-side by topic keywords or date
Pattern 4: Evidence Retrieval
snippet_search(claim_query, limit=5)to find passages making or supporting a claim- To ground a claim within specific papers, pass
paper_ids="<id1>,<id2>,…"(≤100) so snippets come only from that set - For each hit,
get_paper("CorpusId:<corpusId>")for full metadata
Output & Interaction Rules
- Always report which tool was used and the returned row count (no tool exposes a total).
- Present up to 10 results as a table (title, year, venue, citations if fetched), then details for the most relevant few.
- If the user writes in Chinese, present summaries in Chinese; keep titles in their original language.
- After results, offer: Details / Refine / Citations / Snippet / Export / Done.
Handling Asta Responses
Critical Rules
- Prefer batched intent over ping-pong. If a question needs two independent lookups, issue them as parallel MCP calls in one turn, not sequentially.
- Never guess IDs. A fuzzy title goes through
search_paper_by_titlefirst, and ids always carry their prefix. - Check
isErrorbefore reading content, since transport-level status is always 200. - Respect rate limits. An API key buys higher limits but not unlimited ones; stop expanding citation graphs beyond what the user asked for.
- Do not fabricate fields. If Asta returns a null
abstract,venueorcitationCount, say so rather than inventing one.

