
Ntsb Investigations
io.github.pipeworx-iov0.1.0Updated Oct 8, 2026
NTSB Investigations — aviation accident/incident investigations by
Overview
Searches NTSB aviation accident and incident investigations by aircraft registration, make, or model and returns case details.
- What it does
- Exposes one main tool, ntsb_search_investigations, which looks up NTSB aviation accident and incident records by aircraft registration (tail number), make, or model, optionally narrowed by event year or state. Each match returns the NTSB number, event date, location, injury level, phase and occurrence, probable cause, findings, and a docket link. Results include a found flag, data_as_of date, source, count, and the query echoed back.
- When to use it
- Useful when an assistant needs factual aviation accident history for a specific tail number, aircraft make or model, or a year range, such as research, journalism, or due diligence on an aircraft. Not intended for live flight tracking or non-aviation NTSB records.
- Requirements
- Runs as a remote streamable HTTP endpoint at the provider's gateway; no account, API key, or environment variable is declared. A local stdio option is also described, which would require Node.js and npx. Network access to the gateway is needed.
Installation
In SourceWeft
- Open Ntsb Investigations in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.
Other MCP clients
Add this to your client's mcpServers config.
{
"mcpServers": {
"ntsb-investigations": {
"type": "http",
"url": "https://gateway.pipeworx.io/ntsb-investigations/mcp"
}
}
}README
@pipeworx/ntsb-investigations
NTSB aviation accident/incident investigations, searchable by aircraft registration (tail number) or by make/model. Returns NTSB number, event date, location, injury level, phase/occurrence, probable cause, findings, and the NTSB docket link for each match.
Part of Pipeworx — an MCP gateway connecting AI agents to 1721+ live data sources. This is an independent, unofficial integration — not affiliated with, endorsed by, or published by the upstream provider.
Tools
ntsb_search_investigations({ registration?, make?, model?, year?, state?, limit? })— requiresregistrationormake.modelnarrows within a make (matches "172", "172S", "172N", ... by prefix).yearfilters by event year, and is also how a registration match is narrowed when an N-number has been reassigned to a different aircraft over time (see below). Returns{ found, data_as_of, source, count, query, results[] }, or{ found: false, data_as_of, source, query }with no matches.
Auth
Keyless. No upstream account, no API key.
Data sources
- https://data.ntsb.gov/avdata — the NTSB's own eADMS aviation accident
database (
avall.zip, a Microsoft Access.mdbbulk file, confirmed live 2026-10-07 at ~96 MB zipped / ~561 MB unzipped). This is the full, current dataset — the SAME data CAROL's aviation search queries, published as the bulk file NTSB itself ships for exactly this use. - https://data.ntsb.gov/Docket/?NTSBNumber=NNN — the NTSB docket page for
a given
ntsb_number, where the final report and supporting documents are published. Confirmed live (HTTP 200) for a real case pulled from this same dataset.
Why not CAROL's own search, or the newer developer API
CAROL (data.ntsb.gov/carol-main-public), NTSB's modern search UI, was
probed first. Its JSON endpoints (api/Query/Main, api/Query/FileExport)
are undocumented, and every QueryGroups/QueryRules request shape
reverse-engineered from a public third-party CAROL proxy and from NTSB's own
CAROL-Guide.pdf returned HTTP 500 ("An unknown exception occured." / "An
error has occurred.") — the frontend ships no devtools-free schema and the
server gives no hint why the shape it wants differs from what a real browser
session sends (likely a server-side session/anti-forgery requirement CAROL's
own JS sets up before the first query).
NTSB also runs a newer REST API at developer.ntsb.gov (Aviation / Safety
Recommendations / Data Dictionary APIs, via an Azure API Management gateway).
That portal requires signing up for a developer account and a subscription
key before any call succeeds — an authentication wall, which per this repo's
standing rule is never a free pass to build around without a separate key
decision, unlike the plain public bulk file this pack actually reads.
Refreshing the data
node mcps/ntsb-investigations/scripts/refresh.mjs [--local-mdb <path>] [--dry-run] [--r2-mode auto|api|wrangler] [--shard-max-bytes <n>]
Downloads avall.zip, unzips it, runs mdb-export (mdbtools — brew install mdbtools / apt install mdbtools) on five tables (aircraft, events,
narratives, Findings, Events_Sequence, eADMSPUB_DataDictionary), joins
them into one flat case-per-aircraft list, then shards it into the shared
pipeworx-datasets R2 bucket (same bucket as mcps/indiana-code,
mcps/illinois-code, mcps/crs-reports — see scripts/lib/statute-store.mjs)
rather than writing one object:
ntsb/by-make/<slug>.json— one array per manufacturer. A make whose flat shard would exceed ~4 MB (only CESSNA, at this pull) is split further intontsb/by-make/<slug>/<year>.json.- Manufacturers with under 5 rows (most of the ~4,174 distinct raw
acft_makestrings — OCR/data-entry variants, not real distinct makers) are pooled into 16ntsb/by-make/_other/<bucket>.jsonshards by a stable hash, rather than one tiny file each. ntsb/by-registration.json— normalised registration -> the(make, year)pairs it appears under, so a registration query knows which shard(s) to read without a full scan.ntsb/manifest.json—data_as_of,source,record_count, and for every make the exact shard shape (flat/by_year+ its years /bucket), so the read path never has to probe R2 to find out which key a make lives in.
This replaced an earlier single-object design (ntsb/aviation-cases.json,
~27 MB) that re-fetched and parsed the whole dataset on every call with no
cache — two concurrent calls held two full parses in one isolate at once.
src/index.ts now reads the manifest plus exactly the shard(s) one query
needs (one, for a make/model query; the few a reused registration's distinct
(make, year) pairs resolve to, for a registration query) and never reads
the old monolith, which this script deletes from the bucket once the shards
are up. No module-scope cache, on purpose — see the header comment in
src/index.ts. No Supabase migration; storage is R2 only. --local-mdb
re-uses an already-downloaded avall.mdb instead of re-pulling 96 MB.
data_as_of in every response is the date the refresh last ran, not a
per-record field from the upstream.
Registration (N-number) reuse
Tail numbers are reassigned by the FAA to different aircraft over time.
ntsb_search_investigations({ registration }) does not silently collapse
this: every historical match for that tail number comes back, each with its
own event_date, make, and model. Pass year to narrow to one
incarnation when more than one comes back for the same registration.
A broad make may ask you to narrow it
A handful of common manufacturers (CESSNA, at this pull) have enough rows
that their data is split into one shard per year. A query for one of those
makes with neither model nor year would have to read every one of that
make's year-shards to answer — not a full dataset scan, but not the one or
two small objects every other query costs either. Past a small fixed number
of year-shards, the tool refuses with a message asking for year (or
model, which does not change which shards get read but narrows the point
of asking) instead of silently doing the larger read. A make/model/year
combination, or an uncommon make on its own, is unaffected.
Quick Start
Add to your MCP client (Claude Desktop, Cursor, Windsurf, etc.):
What this endpoint actually serves
tools/list at https://gateway.pipeworx.io/ntsb-investigations/mcp returns the tools in the table
above plus the shared Pipeworx meta-tools — ask_pipeworx,
discover_tools, search_within, remember/recall and the rest of the
gateway-wide set. So the tool count you see is larger than this table: a
single-pack endpoint currently lists roughly 30 shared tools alongside the
pack's own. The connection's initialize response states its exact scope, and
is the authoritative answer for a given day.
This is deliberate, not multiplexing by accident. The meta-tools are what let a
scoped connection answer a question this pack does not cover — via
ask_pipeworx, which routes across the whole catalog — without you adding a
second MCP server. There is currently no way to mount a pack endpoint without
them; if the extra schemas cost you more context than the routing is worth,
connect to the full gateway once rather than to several pack endpoints.
Or connect to the full Pipeworx gateway to get every pack's tools listed directly, instead of just this one's:
Both URLs reach the same gateway and the same 1721+ data sources. The
only difference is which pack's tools are listed directly; ask_pipeworx
reaches all of them from either one.
No MCP client? Call it over HTTP
No account needed for the first calls. Inspect any tool: GET https://gateway.pipeworx.io/v1/tools/ntsb_search_investigations. Find one: POST https://gateway.pipeworx.io/v1/tools/search_packs with {"query":"..."}.
Standalone (no gateway account)
This package also runs as a local stdio MCP server — no Pipeworx account, no gateway round-trip:
Or run it directly to confirm it starts:
It speaks MCP over stdin/stdout and answers initialize/tools/list/tools/call
for only this pack's tools — none of the shared meta-tools the gateway
connection above adds. Same source, same tools, no ask_pipeworx routing.
Using with ask_pipeworx
Instead of calling tools directly, you can ask questions in plain English — this works on the pack endpoint above as well as on the full gateway:
The gateway picks the right tool and fills the arguments automatically.
More
License
MIT
Source: README.md at commit bb5ea67
Tools
0Version history
1- v0.1.0LatestOct 8, 2026
