
HKEx Filings
io.github.simonmak-ascentv2.4.1更新於 Oct 2, 2026
Live HKEx (Hong Kong Stock Exchange) regulatory filings for AI agents.
概覽
唯讀的香港交易所監管公告 MCP 伺服器,讓助理搜尋公告並擷取文件內文與表格。
- 功能
- 提供四個唯讀工具:get_server_info、search_filings(時間範圍最多 31 天,可依股票代碼、標題、文件類型、類別與股票名稱篩選)、list_filing_facets(瀏覽某段時間內有哪些公告),以及 get_filing(下載單一文件並擷取其內文與表格)。同一專案也提供本機管線,可將 25 年以上的港交所公告抓取到一個或多個資料庫(共九種),其 stdio MCP 伺服器則在已儲存的語料上提供更完整的工具集。
- 適用情境
- 適合助理需要取得最新港交所監管公告的情境,例如彙整中期報告或追蹤某檔股票的公告。也適合自建 1999 年 4 月以來公告的本機研究、法遵或 RAG 語料,而不必自行撰寫抓取程式。
- 執行需求
- stdio 伺服器透過 uvx 從 PyPI 套件 hkex-filing-scraper 以本機程序執行;本機管線需要 Python、pip install,並設定 DATABASE_TARGET(例如 sqlite 搭配 SQLITE_PATH)。另提供不需 API 金鑰的託管遠端閘道。清單中未宣告任何驗證、環境變數或標頭。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 HKEx Filings,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
HKEx Filing Scraper
[HKEx Filing Scraper — one scraper, many databases]
[CI] [GitHub Release] [PyPI] [License: MIT] [Python 3.10+] [Docs] [MCP] [Glama MCP] [Ruff]
An open-source scraper for 25+ years of HKEx regulatory filings — into any of nine databases, with full-text extraction, graph linking, and a read-only MCP server for AI agents.
The US has EDGAR full-text search. Japan has EDINET. Hong Kong has a search form that returns one page at a time. There is no bulk, machine-readable, full-text corpus of HKEx filings. This builds one.
An open-source Python tool that scrapes 25+ years of Hong Kong Stock Exchange (HKEx) regulatory filings and ingests them into any combination of nine databases — with full-text and table extraction, chunk-level coverage, optional graph linking, and a read-only MCP server so AI agents can query the corpus or the live site.
It speaks the undocumented HKEx JSON API directly, which is faster and more resilient than driving a browser.
Vendors & Integrations
Databases — nine first-class destinations, in documented popularity order (see the support matrix):
- PostgreSQL — production-grade open-source relational
- MySQL / MariaDB — GPL relational servers, one driver
- SQLite — zero-server file database, no install needed
- MongoDB — document database
- Neo4j — property-graph database
- ClickHouse — columnar analytics engine
- DuckDB — in-process analytical engine
- SurrealDB — multi-model graph + document database
AI clients — any MCP-capable agent; ready-made configuration for Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode, Manus, and Perplexity.
Available on — PyPI · Glama · MCP Registry · hosted gateway.
Two Ways to Use It
[Example: install, scrape filings into SQLite, then query the hosted MCP gateway from an AI agent]
Use the Hosted MCP Gateway
POST, Streamable HTTP, no API key:
Four read-only tools: get_server_info, search_filings (a window of at most 31 days, with
optional stock-code, title, document-type, category, and stock-name filters),
list_filing_facets (browse what a window contains), and get_filing (downloads one
document and extracts its text and tables).
[Two ways to reach HKEx filings from an AI agent: the hosted MCP gateway or the local stdio server]
Point a client at it — for example opencode:
Then ask:
Ready-made configuration for Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode,
Manus, and Perplexity is in AI agent support — and for a stored corpus,
the stdio MCP server exposes a wider tool catalog and is published on
Glama. The gateway is listed
in the official MCP Registry as
io.github.simonmak-ascent/hkex-filings.
Featured on Glama — the read-only stdio MCP server is also published on Glama, where Glama scans the built server and scores tool-definition quality (currently 4.7/5).
Quick Start (Local)
Optional extras: excel, postgres, mysql, duckdb, mongodb, clickhouse, neo4j,
mcp, pdf, all, dev.
DATABASE_TARGET is an ordered, comma-separated list of sink ids; the order decides which
sink serves reads. To start with no server:
hkex-scraper runs the full pipeline (metadata + documents + graph); hkex-scraper --full-history covers everything since April 1999. The schema is created automatically.
Full install options and per-sink settings are in Getting started.
Database Support
Every sink is a first-class destination; rows are in documented popularity order. The full matrix — licenses, capability differences, per-engine notes — is in Database sinks.
¹ Source-available, not OSI-approved — labeled exceptions per ADR 0003.
Valid sink ids, in documented order: postgres, mysql, sqlite, mongodb, mariadb, neo4j, clickhouse, duckdb, surrealdb. Set one variable and the same run feeds every sink:
How This Compares
Four ways to get HKEx filings, and what each one costs you.
If you need licensed, redistributable, SLA-backed data, buy the feed. If you need a complete local corpus for research, compliance, or RAG, this replaces the pipeline you would otherwise write yourself.
How It Works
- Phase 1 scrapes filing metadata through a JSF session, splitting the range into monthly
chunks and deduplicating on a 16-character MD5
filingId. - Phase 2 downloads each filing's PDF/HTML/Excel document, extracts text and tables to Markdown, and writes the payload.
- Graph linking (optional) writes
has_filingandreferences_filingedges whenCOMPANY_TABLEis set. - Failure isolation — a failure on one sink is logged and counted but never blocks another; the run exits non-zero if any configured sink failed.
Deeper detail: Architecture · ADR 0002.
Features
- Fast API scraping — direct HKEx JSON API; no browser or Selenium.
- Full history — every filing from April 1999 to today, with chunk-level coverage checks.
- Document processing — PDF/HTML/Excel text and structured tables, extracted to Markdown.
- Multi-sink — any ordered combination of nine databases, each with native idempotent upserts.
- AI-ready — a hosted live MCP gateway plus a local stdio MCP server.
- Resumable and observable — batching, parallel downloads, stalled-job detection, per-sink
counters, and
--coverage-report/--parity-report/--verify. - Optional dependencies — the core is
requests+beautifulsoup4; drivers and document extraction are extras with graceful fallbacks.
Documentation
- Getting started · Configuration · CLI
- Database sinks (matrix) — PostgreSQL, MySQL/MariaDB, SQLite, MongoDB, Neo4j, ClickHouse, DuckDB, SurrealDB
- Live MCP gateway · AI agent support · MCP server
- Architecture · Troubleshooting · Testing
- Roadmap · De-risking register · Upgrading
- What's new · Releasing · Legal & Terms of Use · Changelog
- Docs site: https://hkex-listco-updates.ascent-partners.com/ · Try it locally (
examples/)
Development
Tests are pure unit tests; SQLite and DuckDB contract tests run in-process, and integration tests that need a server are skipped unless that sink is configured. See Testing.
Contributing
See CONTRIBUTING.md; report security issues per SECURITY.md. Ideas and questions are welcome in Discussions.
Built by Ascent Partners.
If this saves you time, a ⭐ on GitHub helps others find it.
Use with Context7
Up-to-date HKEx Filing Scraper documentation is indexed on Context7, so coding agents can pull it into context on demand. With the Context7 MCP server or ctx7 CLI installed, name the library in your prompt:
License
MIT — see LICENSE. That covers this project's code only; optional dependencies
carry their own licenses, notably the pdf extra (PyMuPDF / pymupdf4llm), which is
AGPL-3.0 and deliberately excluded from .[all]. See
docs/legal.md.
Data & Terms of Use: this is a research tool for the undocumented HKEx JSON API, and it is not affiliated with or endorsed by HKEx. Commercial redistribution of HKEx data may require a licensed HKEx feed; see docs/legal.md.
來源:README.md,提交 ce6aed0
工具
0版本歷史
1- v2.4.1最新Oct 2, 2026

