Aidp Profiling Tables

作者 oracle-samples90b42d6c24d4無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Profile an AIDP table — row count, per-column null %, distinct count, min/max/mean, and top-K values. Use when the user asks to profile a table, wants column statistics or a data-quality snapshot, or needs to understand a dataset's shape before using it. Runs bounded Spark SQL via the bundled aidp_sql.py helper.

僅含說明Data & Analytics
AI 產生的概覽

使用 Spark SQL 剖析 AIDP 資料表,產出各欄位的空值率、相異計數、範圍與常見值。

功能
透過隨附的輔助指令碼執行有界 Spark SQL,對單一 AIDP 資料表產出欄位層級的剖析結果。輸出包含資料列數、各欄位空值百分比、近似相異計數、數值欄位的最小值/最大值/平均值、日期範圍,以及類別欄位的前 K 個值。輔助指令碼以 JSON 回傳結果,最後整理成逐欄呈現的表格,並在取樣時加以標註。
適用情境
適用於使用者要求剖析某張資料表、需要欄位統計或資料品質快照,或在使用資料集前需要了解其結構的情況。它著重單一資料表的剖析,而非多資料表分析。
執行需求
需要 OCI 存取權,使用 api_key DEFAULT 設定檔(或僅限工作階段權杖的設定檔),並提供區域、資料湖 OCID、工作區與叢集鍵,以及隨附的 scripts/aidp_sql.py 輔助指令碼及其引用的文件。控制平面查詢使用 oci raw-request,不需要 AIDP MCP 伺服器。需要可連線至 OCI 的網路。

aidp-profiling-tables — single-table profile

Produce a column-level profile of an AIDP table via Spark SQL. Self-contained: control-plane lookups use oci raw-request; profiling SQL runs through the bundled scripts/aidp_sql.py helper. No aidp MCP server is required.

When to use

  • "Profile <table>", "what does <table> look like", "column stats / data quality snapshot".

Workflow

  1. Resolve the table (aidp-catalog-explore / .aidp/catalog.md) → fully-qualified catalog.schema.table and its columns/types. Without a cache, list via oci raw-request: GET /tables?catalogKey=<cat>&schemaKey=<cat.schema> and filter for the table client-side (see references/no-mcp-rest-map.md). Use the column types to pick the right per-column profiling SQL.
  2. Run bounded profiling SQL via the helper (one cell per call; the scratch notebook + kernel are managed for you):
    bash
    python "$PLUGIN_DIR/scripts/aidp_sql.py" --region <r> --datalake <ocid> --workspace <ws> --cluster <key> \  --code "spark.sql('''<profiling SQL>''').show(50, truncate=False)"
    • Overview: SELECT COUNT(*) FROM t (flag if LARGE; sample for the rest).
    • Numeric cols: MIN, MAX, AVG, COUNT, null %, approx distinct (approx_count_distinct).
    • String/categorical: null %, approx_count_distinct, top-K via GROUP BY … ORDER BY count DESC LIMIT k.
    • Date/timestamp: MIN/MAX range, null %. Use TABLESAMPLE/LIMIT on large tables to stay cheap; say when you sampled. The helper returns JSON (status, outputs, spark_job_ids) — parse outputs for the result rows.
  3. Present a per-column table: type, null %, distinct, min/max/mean (numeric), top values (categorical).
  4. Offer to feed findings into .aidp/catalog.md value dictionaries (aidp-catalog-init) and to add data-quality rules (aidp-data-quality).

Reliability rules

  • Profile from real query output, not assumptions; note sampling.
  • For very large tables, profile a sample and label it clearly.
  • The helper mints a UPST from the api_key DEFAULT profile and auto-creates a scratch notebook; pass --session-profile AIDP_SESSION only if your tenancy is session-token-only. On a kernel/auth error, refresh (oci session refresh --profile AIDP_SESSION) and retry.

References

  • references/oci-raw-request.md · references/no-mcp-rest-map.md · pairs with aidp-data-quality, aidp-catalog-init

來源與署名

來源:oracle-samples/oracle-aidp-samples位於ai/claude-code-plugins/oracle-ai-data-platform-workbench-engineer-agent/skills/aidp-profiling-tables提交90b42d6

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 oracle-samples/oracle-aidp-samples 的技能

Aidp Workspace Admin

oracle-samples

Provision and inspect AIDP DataLake instances and workspaces, including private-network workspaces attached to a customer VCN/subnet. Use when the user wants to create/list/get a workspace or DataLake instance, set up a new (e.g. private) AIDP environment, or replicate a customer setup. Create/delete are guarded — confirm before any provisioning.

待分類2026年10月8日

Aidp Volumes

oracle-samples

Work with AIDP volumes — list volumes, browse files inside a volume, upload/download via the PAR flow, and create directories. Use when the user mentions volumes, needs to stage large/binary files, or move data in/out of a volume (distinct from the workspace filesystem). Control-plane via the official `aidp` CLI.

待分類2026年10月8日

Aidp Verified Queries

oracle-samples

維護經過驗證的問題到 Spark SQL 配對庫,讓代理在產生新 SQL 前優先重用可信查詢。

Data & Analytics2026年10月8日

Aidp User Settings

oracle-samples

透過 aidp CLI 或 oci raw-request 備援方式管理 AIDP DataLake 使用者設定與偏好。

Productivity & Workflow2026年10月8日

Aidp Spark Optimization

oracle-samples

指導 Apache Spark 3.5.0 效能調校:分割區、shuffle、join、資料傾斜、記憶體、檔案配置、AQE 與 Delta Lake。

Data & Analytics2026年10月8日

Aidp Semantic Model

oracle-samples

維護 .aidp/semantic.md 業務語意層,定義指標、連接、同義詞與值字典,為自然語言轉 SQL 提供依據。

Data & Analytics2026年10月8日