HKEx Filings

io.github.simonmak-ascentv2.4.1更新于 Oct 2, 2026

Live HKEx (Hong Kong Stock Exchange) regulatory filings for AI agents.

概览

AI 生成的概览

只读的香港交易所监管公告 MCP 服务器,让助手搜索公告并提取文件正文与表格。

功能
提供四个只读工具:get_server_info、search_filings(时间窗口最多 31 天,可按股票代码、标题、文件类型、类别和股票名称筛选)、list_filing_facets(查看某时间窗口内包含哪些公告)以及 get_filing(下载单份文件并提取其正文和表格)。同一项目还提供本地流水线,可将 25 年以上的港交所公告抓取到一个或多个数据库(共九种),其 stdio MCP 服务器在已存储语料上提供更完整的工具集。
适用场景
适合助手需要获取最新港交所监管公告的场景,例如汇总中期报告或跟踪某只股票的公告。也适合自建 1999 年 4 月以来公告的本地研究、合规或 RAG 语料,而不必自己编写抓取程序。
运行要求
stdio 服务器通过 uvx 从 PyPI 包 hkex-filing-scraper 作为本地进程运行;本地流水线需要 Python、pip install,并设置 DATABASE_TARGET(例如 sqlite 配合 SQLITE_PATH)。另提供无需 API 密钥的托管远程网关。清单中未声明任何认证、环境变量或请求头。
安装前请注意
MCP 工具被描述为只读,但本地流水线会把抓取的公告和提取的正文写入你配置的数据库,因此应把 DATABASE_TARGET 指向愿意被写入的目标。POSTGRES_DSN 等数据库连接串包含凭据,不应放入共享配置。抓取依赖未公开的港交所 JSON API,README 说明商业再分发受限并指向其法律文档。

安装

在 SourceWeft 中

  1. 打开 控制台中的 HKEx Filings,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

HKEx Filing Scraper

[HKEx Filing Scraper — one scraper, many databases]

[CI] [GitHub Release] [PyPI] [License: MIT] [Python 3.10+] [Docs] [MCP] [Glama MCP] [Ruff]

An open-source scraper for 25+ years of HKEx regulatory filings — into any of nine databases, with full-text extraction, graph linking, and a read-only MCP server for AI agents.

The US has EDGAR full-text search. Japan has EDINET. Hong Kong has a search form that returns one page at a time. There is no bulk, machine-readable, full-text corpus of HKEx filings. This builds one.

An open-source Python tool that scrapes 25+ years of Hong Kong Stock Exchange (HKEx) regulatory filings and ingests them into any combination of nine databases — with full-text and table extraction, chunk-level coverage, optional graph linking, and a read-only MCP server so AI agents can query the corpus or the live site.

It speaks the undocumented HKEx JSON API directly, which is faster and more resilient than driving a browser.

Vendors & Integrations

Databases — nine first-class destinations, in documented popularity order (see the support matrix):

  • PostgreSQL — production-grade open-source relational
  • MySQL / MariaDB — GPL relational servers, one driver
  • SQLite — zero-server file database, no install needed
  • MongoDB — document database
  • Neo4j — property-graph database
  • ClickHouse — columnar analytics engine
  • DuckDB — in-process analytical engine
  • SurrealDB — multi-model graph + document database

AI clients — any MCP-capable agent; ready-made configuration for Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode, Manus, and Perplexity.

Available on — PyPI · Glama · MCP Registry · hosted gateway.

Two Ways to Use It

Hosted MCP gatewayLocal pipeline
WhatA public endpoint you point an AI agent atThe hkex-scraper CLI
SetupNone — paste a URLpip install + one environment variable
DataLive from HKEx, nothing storedStored in your database(s)
DocsLive MCP gateway · AI agent supportGetting started

[Example: install, scrape filings into SQLite, then query the hosted MCP gateway from an AI agent]

Use the Hosted MCP Gateway

POST, Streamable HTTP, no API key:

text
https://hkex-listco-updates.ascent-partners.com/api/mcp

Four read-only tools: get_server_info, search_filings (a window of at most 31 days, with optional stock-code, title, document-type, category, and stock-name filters), list_filing_facets (browse what a window contains), and get_filing (downloads one document and extracts its text and tables).

[Two ways to reach HKEx filings from an AI agent: the hosted MCP gateway or the local stdio server]

Point a client at it — for example opencode:

json
{  "$schema": "https://opencode.ai/config.json",  "mcp": {    "hkex-live": {      "type": "remote",      "url": "https://hkex-listco-updates.ascent-partners.com/api/mcp"    }  }}

Then ask:

text
Use hkex-live to list the filings published between 2026-09-01 and 2026-09-18,then summarise the interim report.

Ready-made configuration for Claude, ChatGPT, Cursor, VS Code/Copilot, Gemini CLI, opencode, Manus, and Perplexity is in AI agent support — and for a stored corpus, the stdio MCP server exposes a wider tool catalog and is published on Glama. The gateway is listed in the official MCP Registry as io.github.simonmak-ascent/hkex-filings.

Featured on Glama — the read-only stdio MCP server is also published on Glama, where Glama scans the built server and scores tool-definition quality (currently 4.7/5).

Quick Start (Local)

bash
pip install hkex-filing-scraper        # core; SQLite needs no serverpip install "hkex-filing-scraper[all]" # Excel + dotenv + every driver + the MCP servercp .env.example .env                   # then set DATABASE_TARGET (below)hkex-scraper --metadata-only --limit 100

Optional extras: excel, postgres, mysql, duckdb, mongodb, clickhouse, neo4j, mcp, pdf, all, dev.

DATABASE_TARGET is an ordered, comma-separated list of sink ids; the order decides which sink serves reads. To start with no server:

ini
DATABASE_TARGET=sqliteSQLITE_PATH=hkex.db

hkex-scraper runs the full pipeline (metadata + documents + graph); hkex-scraper --full-history covers everything since April 1999. The schema is created automatically. Full install options and per-sink settings are in Getting started.

Database Support

Every sink is a first-class destination; rows are in documented popularity order. The full matrix — licenses, capability differences, per-engine notes — is in Database sinks.

SinkModelLicenseExtraIdempotent upsert
postgresrelationalPostgreSQL LicensepostgresON CONFLICT DO UPDATE
mysql / mariadbrelationalGPLv2mysqlON DUPLICATE KEY UPDATE
sqliterelationalPublic domain—ON CONFLICT DO UPDATE
mongodbdocumentSSPL¹mongodbupdate_one(upsert=True)
neo4jgraphGPLv3 (Community)neo4jMERGE
clickhousecolumnarApache-2.0clickhouseReplacingMergeTree + read-merge
duckdbrelationalMITduckdbON CONFLICT DO UPDATE
surrealdbgraph + documentBSL 1.1¹—UPSERT / RELATE

¹ Source-available, not OSI-approved — labeled exceptions per ADR 0003.

Valid sink ids, in documented order: postgres, mysql, sqlite, mongodb, mariadb, neo4j, clickhouse, duckdb, surrealdb. Set one variable and the same run feeds every sink:

ini
# Order sets read precedence.DATABASE_TARGET=postgres,sqlitePOSTGRES_DSN=postgresql://user:password@localhost:5432/hkexSQLITE_PATH=hkex.db

How This Compares

Four ways to get HKEx filings, and what each one costs you.

This projectHKEXnews web searchBrowser automation you writeLicensed HKEx feed
Bulk exportYesNo — page-at-a-timeYesYes
History to April 1999YesYes, manuallyDepends on your codeYes
Full text of documentsExtracted from PDF/HTML/ExcelNo — you open each fileYou build the extractorVaries by contract
Structured tablesExtracted to MarkdownNoYou build itVaries
Coverage verificationPer-chunk, auditableNot applicableYou build itVendor SLA
Lands in your engine9 engines, any combinationNoWhatever you wire upUsually one format
SpeedJSON API, no browserManualSlower — renders pagesFast
CostFree, MITFreeYour timeSubscription
Commercial redistributionSee docs/legal.mdRestrictedRestrictedLicensed

If you need licensed, redistributable, SLA-backed data, buy the feed. If you need a complete local corpus for research, compliance, or RAG, this replaces the pipeline you would otherwise write yourself.

How It Works

mermaid
flowchart LR    A[HKEx JSON API] --> B[Phase 1: metadata]    B --> C[Canonical record]    C --> D{DATABASE_TARGET}    D --> E[(PostgreSQL)]    D --> F[(MySQL / MariaDB)]    D --> G[(SQLite)]    D --> H[(MongoDB)]    D --> I[(Neo4j)]    D --> J[(ClickHouse)]    D --> K[(DuckDB)]    D --> L[(SurrealDB)]    B --> M[Graph linking]    M --> D    B --> N[Phase 2: download and extract]    N --> C
  • Phase 1 scrapes filing metadata through a JSF session, splitting the range into monthly chunks and deduplicating on a 16-character MD5 filingId.
  • Phase 2 downloads each filing's PDF/HTML/Excel document, extracts text and tables to Markdown, and writes the payload.
  • Graph linking (optional) writes has_filing and references_filing edges when COMPANY_TABLE is set.
  • Failure isolation — a failure on one sink is logged and counted but never blocks another; the run exits non-zero if any configured sink failed.

Deeper detail: Architecture · ADR 0002.

Features

  • Fast API scraping — direct HKEx JSON API; no browser or Selenium.
  • Full history — every filing from April 1999 to today, with chunk-level coverage checks.
  • Document processing — PDF/HTML/Excel text and structured tables, extracted to Markdown.
  • Multi-sink — any ordered combination of nine databases, each with native idempotent upserts.
  • AI-ready — a hosted live MCP gateway plus a local stdio MCP server.
  • Resumable and observable — batching, parallel downloads, stalled-job detection, per-sink counters, and --coverage-report / --parity-report / --verify.
  • Optional dependencies — the core is requests + beautifulsoup4; drivers and document extraction are extras with graceful fallbacks.

Documentation

Development

bash
pip install -e ".[dev,all]"ruff check           # lint (py310, line-length 100)ruff format --check  # formattingpytest               # unit tests (no DB or network required)

Tests are pure unit tests; SQLite and DuckDB contract tests run in-process, and integration tests that need a server are skipped unless that sink is configured. See Testing.

Contributing

See CONTRIBUTING.md; report security issues per SECURITY.md. Ideas and questions are welcome in Discussions.

Built by Ascent Partners.

If this saves you time, a ⭐ on GitHub helps others find it.

Use with Context7

Up-to-date HKEx Filing Scraper documentation is indexed on Context7, so coding agents can pull it into context on demand. With the Context7 MCP server or ctx7 CLI installed, name the library in your prompt:

text
use library /simonmak-ascent/hkex-filing-scraper for API and docs

License

MIT — see LICENSE. That covers this project's code only; optional dependencies carry their own licenses, notably the pdf extra (PyMuPDF / pymupdf4llm), which is AGPL-3.0 and deliberately excluded from .[all]. See docs/legal.md.

Data & Terms of Use: this is a research tool for the undocumented HKEx JSON API, and it is not affiliated with or endorsed by HKEx. Commercial redistribution of HKEx data may require a licensed HKEx feed; see docs/legal.md.

来源:README.md,提交 ce6aed0

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v2.4.1最新Oct 2, 2026