Aidp Profiling Tables

作者 oracle-samples90b42d6c24d4无许可证收录于 2026年10月8日更新于 2026年10月8日

Profile an AIDP table — row count, per-column null %, distinct count, min/max/mean, and top-K values. Use when the user asks to profile a table, wants column statistics or a data-quality snapshot, or needs to understand a dataset's shape before using it. Runs bounded Spark SQL via the bundled aidp_sql.py helper.

仅含说明Data & Analytics
AI 生成的概览

使用 Spark SQL 对 AIDP 表进行剖析,生成各列的空值率、去重计数、范围和常见取值。

功能
通过随附的辅助脚本运行有界 Spark SQL,对单张 AIDP 表生成列级剖析结果。输出包括行数、各列空值百分比、近似去重计数、数值列的最小值/最大值/平均值、日期范围,以及分类列的前 K 个取值。辅助脚本以 JSON 返回结果,最终整理为按列展示的表格,并在采样时予以说明。
适用场景
适用于用户要求剖析某张表、需要列统计或数据质量快照,或在使用数据集前需要了解其结构的情况。它面向单表剖析,而非多表分析。
运行要求
需要具备 OCI 访问权限,使用 api_key DEFAULT 配置文件(或仅会话令牌的配置文件),并提供区域、数据湖 OCID、工作区和集群键,以及随附的 scripts/aidp_sql.py 辅助脚本及其引用的文档。控制面查询使用 oci raw-request,无需 AIDP MCP 服务器。需要访问 OCI 的网络连接。

aidp-profiling-tables — single-table profile

Produce a column-level profile of an AIDP table via Spark SQL. Self-contained: control-plane lookups use oci raw-request; profiling SQL runs through the bundled scripts/aidp_sql.py helper. No aidp MCP server is required.

When to use

  • "Profile <table>", "what does <table> look like", "column stats / data quality snapshot".

Workflow

  1. Resolve the table (aidp-catalog-explore / .aidp/catalog.md) → fully-qualified catalog.schema.table and its columns/types. Without a cache, list via oci raw-request: GET /tables?catalogKey=<cat>&schemaKey=<cat.schema> and filter for the table client-side (see references/no-mcp-rest-map.md). Use the column types to pick the right per-column profiling SQL.
  2. Run bounded profiling SQL via the helper (one cell per call; the scratch notebook + kernel are managed for you):
    bash
    python "$PLUGIN_DIR/scripts/aidp_sql.py" --region <r> --datalake <ocid> --workspace <ws> --cluster <key> \  --code "spark.sql('''<profiling SQL>''').show(50, truncate=False)"
    • Overview: SELECT COUNT(*) FROM t (flag if LARGE; sample for the rest).
    • Numeric cols: MIN, MAX, AVG, COUNT, null %, approx distinct (approx_count_distinct).
    • String/categorical: null %, approx_count_distinct, top-K via GROUP BY … ORDER BY count DESC LIMIT k.
    • Date/timestamp: MIN/MAX range, null %. Use TABLESAMPLE/LIMIT on large tables to stay cheap; say when you sampled. The helper returns JSON (status, outputs, spark_job_ids) — parse outputs for the result rows.
  3. Present a per-column table: type, null %, distinct, min/max/mean (numeric), top values (categorical).
  4. Offer to feed findings into .aidp/catalog.md value dictionaries (aidp-catalog-init) and to add data-quality rules (aidp-data-quality).

Reliability rules

  • Profile from real query output, not assumptions; note sampling.
  • For very large tables, profile a sample and label it clearly.
  • The helper mints a UPST from the api_key DEFAULT profile and auto-creates a scratch notebook; pass --session-profile AIDP_SESSION only if your tenancy is session-token-only. On a kernel/auth error, refresh (oci session refresh --profile AIDP_SESSION) and retry.

References

  • references/oci-raw-request.md · references/no-mcp-rest-map.md · pairs with aidp-data-quality, aidp-catalog-init

来源与署名

来源:oracle-samples/oracle-aidp-samples位于ai/claude-code-plugins/oracle-ai-data-platform-workbench-engineer-agent/skills/aidp-profiling-tables提交90b42d6

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 oracle-samples/oracle-aidp-samples 的技能

Aidp Workspace Admin

oracle-samples

Provision and inspect AIDP DataLake instances and workspaces, including private-network workspaces attached to a customer VCN/subnet. Use when the user wants to create/list/get a workspace or DataLake instance, set up a new (e.g. private) AIDP environment, or replicate a customer setup. Create/delete are guarded — confirm before any provisioning.

待分类2026年10月8日

Aidp Volumes

oracle-samples

Work with AIDP volumes — list volumes, browse files inside a volume, upload/download via the PAR flow, and create directories. Use when the user mentions volumes, needs to stage large/binary files, or move data in/out of a volume (distinct from the workspace filesystem). Control-plane via the official `aidp` CLI.

待分类2026年10月8日

Aidp Verified Queries

oracle-samples

维护经过验证的问题到 Spark SQL 配对库,让智能体在生成新 SQL 前优先复用可信查询。

Data & Analytics2026年10月8日

Aidp User Settings

oracle-samples

通过 aidp CLI 或 oci raw-request 备用方式管理 AIDP DataLake 用户设置与偏好。

Productivity & Workflow2026年10月8日

Aidp Spark Optimization

oracle-samples

指导 Apache Spark 3.5.0 性能调优:分区、shuffle、连接、倾斜、内存、文件布局、AQE 与 Delta Lake。

Data & Analytics2026年10月8日

Aidp Semantic Model

oracle-samples

维护 .aidp/semantic.md 业务语义层,定义指标、连接、同义词和值字典,为自然语言转 SQL 提供依据。

Data & Analytics2026年10月8日