Sec Financial Statements

io.github.pipeworx-iov0.1.0更新於 Oct 1, 2026

SEC Financial Statement & Notes data sets — dimensional facts + note text.

已驗證Streamable HTTP可網頁執行DatabasesFinanceData & Analytics

概覽

AI 產生的概覽

查詢 SEC 財務報表與附註資料集,取得帶維度的 XBRL 事實與帶標籤的附註文字,可跨申報人與期間檢索。

功能
透過四個工具存取 SEC DERA 財務報表與附註的大量資料:company_financial_facts 依公司名稱或 CIK 回傳單一申報人的帶維度事實,包含分部、地區與避險指定等軸;tag_cross_company_screen 對某一期間內所有申報人的同一標籤值排序;note_text_search 對帶標籤的附註文字區塊做全文檢索;fsn_coverage 回報已載入的期間,以及各期間包含的提交、事實與附註區塊數量。它涵蓋 companyfacts 與 frames 未提供的帶維度事實。
適用情境
適合跨多家申報人篩選某個 XBRL 標籤、取出分部或地區拆分,或跨申報文件檢索揭露文字。它與申報文字檢索互補:後者用於在單一申報的正文中定位事實,本服務則用於跨公司與期間查詢結構化事實與附註。
執行需求
遠端 streamable HTTP 端點,使用閘道網址;呼叫端不需要帳號或 API 金鑰,資料憑證由閘道注入。本機 stdio 版本透過 npx 執行,需要 Node.js。涵蓋範圍僅限實際已載入的期間,建議先呼叫 fsn_coverage。
安裝前請注意
README 說明在提交時尚未載入任何期間,此套件也尚未部署,因此在資料匯入前查詢可能回傳空結果。涵蓋範圍有限,且僅由 fsn_coverage 回報。標籤是精確、區分大小寫的 XBRL 元素名稱,且沒有股票代號欄位,查詢需使用公司名稱或 CIK。修訂申報不會被合併為重述感知檢視。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 Sec Financial Statements,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。

其他 MCP 客戶端

把它新增到你客戶端的 mcpServers 設定中。

{
  "mcpServers": {
    "sec-financial-statements": {
      "type": "http",
      "url": "https://gateway.pipeworx.io/sec-financial-statements/mcp"
    }
  }
}

README

@pipeworx/sec-financial-statements

Dimensioned XBRL facts and tagged note text from SEC's DERA "Financial Statement and Notes" data sets — the layer beneath data.sec.gov companyfacts/frames (what edgar and sec-xbrl call), which expose undimensioned facts only. A segment-level revenue number, a geography breakdown, a footnote-tagged hedge notional — none of that is reachable through companyfacts/frames at all; it only exists in SEC's bulk data sets. Hosted, keyless to the caller.

Part of Pipeworx — an MCP gateway connecting AI agents to 1686+ live data sources.

Status at commit time: schema + tools + loader are code-complete; NO PERIOD IS LOADED YET. The migration (supabase/migrations/219_sec_financial_statements.sql) has not been applied to production and this pack has not been deployed. See "Why this ships without data" below — this is deliberate, not an oversight, and the next step is written down there.

Tools

  • company_financial_facts(company?, cik?, tag?, limit?) — dimensioned facts for one filer (give a name, ILIKE-matched, or a CIK). Each fact carries its period, value, unit and — when dimensioned — the segment/geography axis it belongs to, resolved through sec_fsn_dimensions rather than left as a bare hash. This is the tool that answers what companyfacts/frames structurally cannot: "Flutter Entertainment's cross-currency interest rate swap notional" is a DerivativeNotionalAmount fact dimensioned by hedge designation, not a consolidated total.
  • tag_cross_company_screen(tag, version?, ddate?, dimensioned_only?, limit?) — every filer's value for ONE tag in one period, ranked — the inverse of a per-company lookup ("who reported the highest AccountsPayableCurrent this quarter").
  • note_text_search(query, company?, tag?, limit?) — full-text search (Postgres english tsvector) over tagged note/footnote text blocks — the disclosure prose behind the numbers. Searches tagged note blocks only, not full filing text; see "edgar_filing_text vs. this pack" below.
  • fsn_coverage() — which periods are loaded and how many submissions/facts/note blocks each holds. Call this first when a lookup returns nothing — coverage is real and finite, never silently assumed.

edgar_filing_text vs. this pack — complementary, not duplicative

edgar_filing_text (fixed 2026-09-29, fleet #2500, commit 7ab72cdca) does pipe-separated multi-phrase prose search over a filing's rendered text and attaches the one XBRL fact matched to that exact accession when a phrase reads as a financial concept. This pack is the opposite shape: a structured database of every tagged numeric fact (dimensioned included) and every tagged note text block across whichever periods are loaded, queryable by tag/company/period without needing to guess a search phrase that happens to appear verbatim in the rendered document. Use edgar_filing_text to find a fact inside one specific filing's prose; use this pack to screen across companies, pull a full dimensional breakdown, or search note text across many filings at once.

Why this ships without data

Sizes were explicitly UNVERIFIED when this pack was filed (fleet #2506) — "measure first." Measured live against the newest available period (2026_08_notes.zip, HTTP 200, content-length: 312,729,503 compressed): ~2.46GB uncompressed across the six tables this pack loads, num.tsv alone 5,563,767 rows / 1.14GB. Loading even one period is a real, multi-GB bulk operation over SEC's shared fleet egress (CLAUDE.md, fleet #1245) — not something to run casually inside a build session. Parsing and row-mapping were validated locally against the real 2026_08_notes.zip (every table's header/columns confirmed, dimensioned vs. undimensioned facts correctly distinguished, note text and submission rows read cleanly) — the loader is not speculative, it has been run against real SEC bytes, just not into production yet.

Two things gate actually loading data and shipping this pack live:

  1. The migration has to land on main and be applied via db-migrate.yml (workflow_dispatch, gh workflow run db-migrate.yml -f file=219_sec_financial_statements.sql — dry-run first with -f dry_run=true), per CLAUDE.md's local-copy rule ("migration lands ALONE on main… before the pack ships"). This build's own fleet task (#2506) explicitly says "Commit, do NOT push — report SHAs to the dispatching PM" — so the push/apply/deploy decision is the dispatching PM's, not this session's, and is reported rather than executed here. (A direct supabase db query --file … --linked apply — the documented ad-hoc testing path — was attempted for local verification and was refused by the session's own permission policy as a protected infrastructure-apply action; that refusal was respected rather than routed around, which is the correct behavior for anyone hitting the same wall.)
  2. The first period load is a real ingest run, not a CI step: node scripts/ingest-sec-fsn.mjs --period 2026_08 (or --list-periods to see what SEC currently publishes). Budget real wall time — num.tsv alone is 5.56M rows at 3,000/batch.

Data + refresh

Source: https://www.sec.gov/data-research/sec-markets-data/financial-statement-notes-data-sets — quarterly zips 2009q1 through 2025q2, then monthly (2026_08, etc.) from 2025_08 onward (SEC changed cadence mid-2025). US federal government data, public domain — no reuse-grant question. Six of the eight TSVs each zip carries are loaded (sub, tag, dim, pre, num, txt); cal (calculation linkbase) and ren (rendering metadata) are not — they serve statement rendering, not the fact/note lookups this pack answers.

  • Loader: node scripts/ingest-sec-fsn.mjs --period <YYYY_MM|YYYYqN> — one period per run, same credential resolution and batched-upsert-with-retry shape as scripts/ingest-sec-13f.mjs. --only sub,tag,dim,pre,num,txt loads a subset; --zip <path> skips the download for a local/already-fetched archive.
  • Re-running an already-loaded period is a safe no-op upsert: num/txt key on a content hash of every loaded column (SEC's own documented natural key for num.tsv is 9 columns, several nullable — a hash collapses exact SEC-side duplicates without losing genuinely distinct rows, same pattern as leie_exclusions / csl_entries / ferc_eqr, migration 211).
  • Backfill plan: load the newest period first (proves the pipeline end-to-end fastest), then step backward one period per run as budget allows. fsn_coverage() always reports exactly what's loaded — no response ever implies coverage beyond that.
  • Schema and indexes: supabase/migrations/219_sec_financial_statements.sql. Row-level security is enabled on all six tables, anon/authenticated revoked (migration 211's ferc_eqr pattern) — only the gateway's own injected data credential can read.

What this cannot tell you

  • Coverage is whatever has actually been loaded, not the full 2009-present history SEC publishes. fsn_coverage() is the source of truth; a lookup outside loaded periods returns nothing and says so rather than reading as "this fact doesn't exist."
  • No ticker column — SEC's DERA data sets key on CIK/accession, not ticker. company_financial_facts takes a company name (ILIKE) or a CIK; an ambiguous name match returns the candidate CIKs instead of silently picking one.
  • Tags are exact, case-sensitive XBRL element names ("Revenues", not "revenue") — this pack does not fuzzy-match tag names, because a near-miss tag is a different accounting concept, not a typo.
  • This is what filers tagged, not a restatement-aware view. An amended filing's facts sit alongside the original's under different accessions; this pack does not collapse them (contrast sec-13f's amendment handling, which has no analog here since XBRL facts are versioned by filing, not by position).

Auth

Keyless to the caller — the data credentials are injected by the gateway, never exposed.

Quick Start

Add to your MCP client (Claude Desktop, Cursor, Windsurf, etc.):

json
{  "mcpServers": {    "sec-financial-statements": {      "url": "https://gateway.pipeworx.io/sec-financial-statements/mcp"    }  }}

What this endpoint actually serves

tools/list at https://gateway.pipeworx.io/sec-financial-statements/mcp returns the tools in the table above plus the shared Pipeworx meta-tools — ask_pipeworx, discover_tools, search_within, remember/recall and the rest of the gateway-wide set. So the tool count you see is larger than this table: a single-pack endpoint currently lists roughly 30 shared tools alongside the pack's own. The connection's initialize response states its exact scope, and is the authoritative answer for a given day.

This is deliberate, not multiplexing by accident. The meta-tools are what let a scoped connection answer a question this pack does not cover — via ask_pipeworx, which routes across the whole catalog — without you adding a second MCP server. There is currently no way to mount a pack endpoint without them; if the extra schemas cost you more context than the routing is worth, connect to the full gateway once rather than to several pack endpoints.

Or connect to the full Pipeworx gateway to get every pack's tools listed directly, instead of just this one's:

json
{  "mcpServers": {    "pipeworx": {      "url": "https://gateway.pipeworx.io/mcp"    }  }}

Both URLs reach the same gateway and the same 1686+ data sources. The only difference is which pack's tools are listed directly; ask_pipeworx reaches all of them from either one.

No MCP client? Call it over HTTP

bash
curl -X POST https://gateway.pipeworx.io/v1/tools/company_financial_facts \  -H 'Content-Type: application/json' \  -d '{"company":"Apple","tag":"Revenues"}'

No account needed for the first calls. Inspect any tool: GET https://gateway.pipeworx.io/v1/tools/company_financial_facts. Find one: POST https://gateway.pipeworx.io/v1/tools/search_packs with {"query":"..."}.

Standalone (no gateway account)

This package also runs as a local stdio MCP server — no Pipeworx account, no gateway round-trip:

json
{  "mcpServers": {    "sec-financial-statements": {      "command": "npx",      "args": ["-y", "@pipeworx/mcp-sec-financial-statements"]    }  }}

Or run it directly to confirm it starts:

bash
npx -y @pipeworx/mcp-sec-financial-statements

It speaks MCP over stdin/stdout and answers initialize/tools/list/tools/call for only this pack's tools — none of the shared meta-tools the gateway connection above adds. Same source, same tools, no ask_pipeworx routing.

Using with ask_pipeworx

Instead of calling tools directly, you can ask questions in plain English — this works on the pack endpoint above as well as on the full gateway:

ask_pipeworx({ question: "your question about Sec Financial Statements data" })

The gateway picks the right tool and fills the arguments automatically.

More

License

MIT

來源:README.md,提交 b2c17c0

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.1.0最新Oct 1, 2026