aidp-catalog-init — build the catalog grounding file
Walk the AIDP catalog tree and generate .aidp/catalog.md — the cached, user-editable grounding file that
makes subsequent NL-to-SQL fast and accurate.
Discovery is pure control-plane — no SQL, no compute (except optional --with-counts, which uses
the bundled SQL helper). Self-contained: no aidp MCP required.
When to use
- First-time setup, or
--refreshafter schema changes.
Engine — official aidp CLI (control-plane, no compute)
Preferred engine is the official Oracle aidp CLI; oci raw-request is the fallback when the CLI isn't
installed. Both hit the same data-plane REST API with the same auth — see
references/aidp-cli-map.md for the full skill→command map and
references/oci-raw-request.md for base URL + auth ladder + conventions.
CLI (preferred):
Fallback (no CLI installed) — oci raw-request (LIVE-VERIFIED 20240831 / dataLakes /
--profile DEFAULT — see references/no-mcp-rest-map.md):
- Single table / columns —
aidp schema get-table(or the RESTtables?…list, which returns columns, types, and properties); filter to the one table client-side by its key (no dedicated single-table param confirmed — see no-mcp-rest-map.md). - Per-endpoint params are required: a bare path returns
400 InvalidParameter: query param X must not be null, which names the missing param. - On
401/403/"Security Token", follow the auth ladder (refreshAIDP_SESSION, retry with--auth security_token) in oci-raw-request.md.
Process
- Walk the tree (no compute):
aidp catalog list→ for each,aidp schema list --catalog-key→ for each,aidp schema list-tables --catalog-key --schema-key(columns, types, properties) — or the REST fallback above. For large catalogs, fan out one subagent per catalog to parallelize discovery. - Capture grounding hints (this is what raises NL-SQL accuracy):
- FK/join hints — infer likely join keys from naming (
*_sk,*_id, shared column names) and any declared keys in the table properties. Record them so the agent doesn't guess joins later. - Value dictionaries — for low-cardinality categorical columns, note canonical values/format
(prevents wrong WHERE literals like "California" vs "CA"). Pull distinct values only when cheap
(
--with-countspath), or mark TODO. - Large-table flags — flag big fact tables ("always filter by date").
- FK/join hints — infer likely join keys from naming (
- Enrich from the codebase if present (existing notebooks, SQL files, CLAUDE.md) for descriptions.
- Write
.aidp/catalog.mdwith sections: Quick Reference (concept→table), Catalogs → schemas → tables (columns, types, join keys, flags), Value dictionaries, Gotchas. Preserve user edits + HTML comments on--refresh; flag removed tables with<!-- REMOVED -->. - Summarize to the user (N catalogs / schemas / tables, large tables flagged) and suggest next steps
(
aidp-semantic-modelfor metrics,aidp-analyzing-datato ask questions).
Options
-
--refresh— regenerate, preserving user edits and Quick-Reference rows. -
--catalog <name>— limit to one catalog. -
--with-counts— also fetch row counts / distinct values via the bundled SQL helper (uses the cluster, off by default — it costs compute and needs a running cluster):Returns JSON with
status/outputs/spark_job_ids; mints a UPST from the api_key DEFAULT profile and auto-creates a scratch notebook (no AIDP_SESSION required). See references/oci-raw-request.md for the control-plane side.
Output format (.aidp/catalog.md)
Notes
- Resolve
<region>/<DATALAKE_OCID>/<workspace>explicitly — catalog calls are scoped to the DataLake; the SQL helper is scoped to a workspace + cluster. .aidp/is git-ignored — it's a per-project cache, not shipped with the plugin.- Auto-Populate Catalog Extractor (bulk auto-cataloging from Object Storage) has a REST surface at
…/dataLakes/<OCID>/extractors(NOT/metadataExtractors, which 404s — an earlier note probed the wrong path). LIVE-VERIFIED 2026-06-12:GET …/20240831/dataLakes/<OCID>/extractors→ 200{"items":[]}. Surface:GET/POST/DELETE /extractors,GET /extractors/<key>/extractedEntities,GET /extractors/<key>/extractedTables/<name>,POST /extractors/<key>/actions/manageExtractedEntities(accept/reject/import), lifecycleACCEPTED→IN_PROGRESS→SUCCEEDED/FAILED/IN_REVIEW. This complements (does not replace) the discovery walk above andaidp-ingest-file-to-table. Probe the create/manage write paths live (need an Object Storage source) before relying on them. - The aidp MCP is an optional accelerator — if one is configured you may use
list_catalogs/list_schemas/list_tables/get_tableinstead of the raw calls, but it is not required.
References
- references/aidp-cli-map.md — skill → official
aidpCLI command map (primary engine) - references/oci-raw-request.md · references/no-mcp-rest-map.md · references/semantic-model.md
