Aidp Ingest File To Table

作者 oracle-samples90b42d6c24d4無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Load a data file (CSV/JSON/Parquet/etc.) into a managed AIDP Delta table. Use when the user wants to ingest a file into a table, create a table from a file, or land raw data in the lakehouse. Supports the 1-step path and the 3-step upload→infer→create path. Control-plane via the official `aidp` CLI.

僅含說明Data & Analytics
AI 產生的概覽

透過 aidp CLI 或 OCI REST 呼叫,把 CSV、JSON 或 Parquet 檔案載入受管理的 AIDP Delta 資料表。

功能
這個技能引導代理把資料檔案落地到 DataLake schema/tables 資源上的受管理 AIDP Delta 資料表。它支援一步式 create-table 路徑,以及三步流程:產生上傳目標、推斷並預覽結構描述、以最終確定的欄位建立資料表。它也涵蓋輪詢非同步建表、驗證結果資料表,以及已記錄的限制,例如僅支援逗號分隔符、外部表不支援多行 JSON。
適用情境
當使用者想把 CSV、JSON、Parquet 或類似檔案載入資料表、從檔案建立資料表,或把原始資料落地到湖倉時使用。它適合檔案到資料表的擷取,而非持續串流擷取或外部資料來源擷取。
執行需求
需要官方 Oracle aidp CLI,並設定 API 金鑰驗證設定檔、區域與 DataLake 執行個體 OCID;未安裝 CLI 時可用 oci raw-request 作為備援。需要存取 AIDP DataLake 控制平面的網路,變更類操作應先與使用者確認。這個技能不附帶指令碼,只有說明文件與引用的參考文件。

aidp-ingest-file-to-table — file → managed Delta table

Land a file into a managed AIDP table, either in one call or via the staged 3-step flow when you need to review/adjust the inferred schema. This is a control-plane flow on the DataLake schema/tables resource. Primary engine: the official Oracle aidp CLI (same REST API + auth); oci raw-request is the fallback when the CLI isn't installed.

When to use

  • "Load this CSV/JSON into a table", "create a table from <file>", "ingest <file> into the lakehouse".

CLI (preferred)

Per references/aidp-cli-map.md: schema generate-temp-file-upload-target → schema infer / infer-with-preview → schema create-data-table / create-table (also schema retrieve-par). All commands take --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region <r>.

bash
# 3-step (control): stage → infer → createaidp schema generate-temp-file-upload-target --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region us-ashburn-1   # returns upload target / PAR (also: retrieve-par)aidp schema infer-with-preview              --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region us-ashburn-1   # review columns/types/preview (or: infer)aidp schema create-data-table --body-file .aidp/payloads/create-data-table-<name>.json \  --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region us-ashburn-1                                            # or: create-table

Mutating ops (create-data-table/create-table, upload): persist the body to .aidp/payloads/ and confirm with the user before running (see references/payloads.md).

Fallback (no CLI) — same REST + auth via oci raw-request against …/20240831/dataLakes/<OCID>/… (auth ladder in references/oci-raw-request.md): POST /tables/actions/uploadDataFile (multipart/binary may need PAR upload — see aidp-volumes), POST /tables/actions/inferSchema, POST /tables/actions/createTable (with catalogKey, schemaKey, table name, finalized columns, source format, load options), verify GET /tables?catalogKey=<cat>&schemaKey=<cat.schema>.

Verify-first (no-fabrication): the upload/infer/create action shapes are UNVERIFIED in this env (not yet in references/rest-endpoint-map.md). Confirm with a live probe (start with a GET /tables?catalogKey=…&schemaKey=… 200 against the target schema) before any write; record results.

Live-verified 2026-06-10 on de-agent (CSV → de_ingest_test, 3 rows) — correction: the uploadDataFile / inferSchema / createTable action names above are WRONG. The working flow is the schema-resource 3-step: (1) generate-temp-file-upload-target returns a PAR + ociFilePath; (2) PUT the file bytes to the PAR (HTTP 200); (3) infer-with-preview — its location MUST be the ociFilePath OCI URI, not the uploadKey (passing uploadKey → 400); (4) create-data-table returns 202 + a datalake-async-operation-key (poll to SUCCEEDED). create-data-table is HEADERLESS/POSITIONAL: header=true is ignored at create, so tableFields must use the reader column names _c0/_c1/_c2… — naming them id/name/amt fails the async op with UNRESOLVED_COLUMN. Rename afterward via ALTER TABLE … RENAME COLUMN.

Workflow

  1. Confirm the source file location (workspace path or volume) and the target catalog.schema.table (create the schema first if needed).
  2. 1-step (simple): aidp schema create-table referencing the source file, format, and options — fastest when the schema infers cleanly.
  3. 3-step (control): generate-temp-file-upload-target → infer-with-preview (review columns/types with the user; fix types/headers/delimiters) → create-data-table with the finalized columns.
  4. Async: table creation may return 202 with an async-operation key — poll until terminal (async convention in references/oci-raw-request.md; track via aidp-observability).
  5. Verify with aidp schema list-tables / GET /tables?…; report the fully-qualified table name and row/column summary.

Gotchas (documented limits, no workaround)

  • Delimited files: comma only — auto-populate "Doesn't support delimiters other than comma" (platform reference §42 Known Issues #15). Pre-convert tab/pipe/semicolon-delimited files to CSV before ingest.
  • No multi-line JSON for external tables — "Can't create external tables with multi-line JSON" (platform reference §42 Known Issues #12). Use newline-delimited JSON (one record per line) for external tables.

Notes

  • Big files: prefer landing into a volume / object storage and loading from there; mind cluster memory.
  • For continuous/streaming or external-source ingestion, use the spark-connectors plugin + aidp-federate, not this skill (this is file→table).
  • Clean up temporary tables created during validation.

References

  • references/aidp-cli-map.md · references/payloads.md · references/oci-raw-request.md · references/rest-endpoint-map.md
  • pairs with aidp-workspace-files, aidp-volumes, aidp-profiling-tables

來源與署名

來源:oracle-samples/oracle-aidp-samples位於ai/claude-code-plugins/oracle-ai-data-platform-workbench-engineer-agent/skills/aidp-ingest-file-to-table提交90b42d6

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 oracle-samples/oracle-aidp-samples 的技能

Aidp Workspace Admin

oracle-samples

Provision and inspect AIDP DataLake instances and workspaces, including private-network workspaces attached to a customer VCN/subnet. Use when the user wants to create/list/get a workspace or DataLake instance, set up a new (e.g. private) AIDP environment, or replicate a customer setup. Create/delete are guarded — confirm before any provisioning.

待分類2026年10月8日

Aidp Volumes

oracle-samples

Work with AIDP volumes — list volumes, browse files inside a volume, upload/download via the PAR flow, and create directories. Use when the user mentions volumes, needs to stage large/binary files, or move data in/out of a volume (distinct from the workspace filesystem). Control-plane via the official `aidp` CLI.

待分類2026年10月8日

Aidp Verified Queries

oracle-samples

維護經過驗證的問題到 Spark SQL 配對庫,讓代理在產生新 SQL 前優先重用可信查詢。

Data & Analytics2026年10月8日

Aidp User Settings

oracle-samples

透過 aidp CLI 或 oci raw-request 備援方式管理 AIDP DataLake 使用者設定與偏好。

Productivity & Workflow2026年10月8日

Aidp Spark Optimization

oracle-samples

指導 Apache Spark 3.5.0 效能調校:分割區、shuffle、join、資料傾斜、記憶體、檔案配置、AQE 與 Delta Lake。

Data & Analytics2026年10月8日

Aidp Semantic Model

oracle-samples

維護 .aidp/semantic.md 業務語意層,定義指標、連接、同義詞與值字典,為自然語言轉 SQL 提供依據。

Data & Analytics2026年10月8日