Aidp Ingest File To Table

by oracle-samples90b42d6c24d4No licenseListed Oct 8, 2026Updated Oct 8, 2026

Load a data file (CSV/JSON/Parquet/etc.) into a managed AIDP Delta table. Use when the user wants to ingest a file into a table, create a table from a file, or land raw data in the lakehouse. Supports the 1-step path and the 3-step upload→infer→create path. Control-plane via the official `aidp` CLI.

Instructions onlyData & Analytics
AI-generated overview

Loads a CSV, JSON, or Parquet file into a managed AIDP Delta table via the aidp CLI or OCI REST calls.

What it does
This skill guides an agent through landing a data file into a managed AIDP Delta table on the DataLake schema/tables resource. It supports a one-step create-table path and a three-step flow of generating an upload target, inferring the schema with a preview, and creating the table with finalized columns. It also covers polling asynchronous table creation, verifying the resulting table, and documented limits such as comma-only delimiters and no multi-line JSON for external tables.
When to use it
Use it when a user wants to load a CSV, JSON, Parquet, or similar file into a table, create a table from a file, or land raw data in the lakehouse. It fits file-to-table ingestion, not continuous streaming or external-source ingestion.
Requirements
Requires the official Oracle aidp CLI with an API-key auth profile, region, and DataLake instance OCID; oci raw-request is a fallback when the CLI is absent. Network access to the AIDP DataLake control plane is needed, and mutating operations should be confirmed with the user. It ships no scripts, only instructions and referenced documents.

aidp-ingest-file-to-table — file → managed Delta table

Land a file into a managed AIDP table, either in one call or via the staged 3-step flow when you need to review/adjust the inferred schema. This is a control-plane flow on the DataLake schema/tables resource. Primary engine: the official Oracle aidp CLI (same REST API + auth); oci raw-request is the fallback when the CLI isn't installed.

When to use

  • "Load this CSV/JSON into a table", "create a table from <file>", "ingest <file> into the lakehouse".

CLI (preferred)

Per references/aidp-cli-map.md: schema generate-temp-file-upload-target → schema infer / infer-with-preview → schema create-data-table / create-table (also schema retrieve-par). All commands take --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region <r>.

bash
# 3-step (control): stage → infer → createaidp schema generate-temp-file-upload-target --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region us-ashburn-1   # returns upload target / PAR (also: retrieve-par)aidp schema infer-with-preview              --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region us-ashburn-1   # review columns/types/preview (or: infer)aidp schema create-data-table --body-file .aidp/payloads/create-data-table-<name>.json \  --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region us-ashburn-1                                            # or: create-table

Mutating ops (create-data-table/create-table, upload): persist the body to .aidp/payloads/ and confirm with the user before running (see references/payloads.md).

Fallback (no CLI) — same REST + auth via oci raw-request against …/20240831/dataLakes/<OCID>/… (auth ladder in references/oci-raw-request.md): POST /tables/actions/uploadDataFile (multipart/binary may need PAR upload — see aidp-volumes), POST /tables/actions/inferSchema, POST /tables/actions/createTable (with catalogKey, schemaKey, table name, finalized columns, source format, load options), verify GET /tables?catalogKey=<cat>&schemaKey=<cat.schema>.

Verify-first (no-fabrication): the upload/infer/create action shapes are UNVERIFIED in this env (not yet in references/rest-endpoint-map.md). Confirm with a live probe (start with a GET /tables?catalogKey=…&schemaKey=… 200 against the target schema) before any write; record results.

Live-verified 2026-06-10 on de-agent (CSV → de_ingest_test, 3 rows) — correction: the uploadDataFile / inferSchema / createTable action names above are WRONG. The working flow is the schema-resource 3-step: (1) generate-temp-file-upload-target returns a PAR + ociFilePath; (2) PUT the file bytes to the PAR (HTTP 200); (3) infer-with-preview — its location MUST be the ociFilePath OCI URI, not the uploadKey (passing uploadKey → 400); (4) create-data-table returns 202 + a datalake-async-operation-key (poll to SUCCEEDED). create-data-table is HEADERLESS/POSITIONAL: header=true is ignored at create, so tableFields must use the reader column names _c0/_c1/_c2… — naming them id/name/amt fails the async op with UNRESOLVED_COLUMN. Rename afterward via ALTER TABLE … RENAME COLUMN.

Workflow

  1. Confirm the source file location (workspace path or volume) and the target catalog.schema.table (create the schema first if needed).
  2. 1-step (simple): aidp schema create-table referencing the source file, format, and options — fastest when the schema infers cleanly.
  3. 3-step (control): generate-temp-file-upload-target → infer-with-preview (review columns/types with the user; fix types/headers/delimiters) → create-data-table with the finalized columns.
  4. Async: table creation may return 202 with an async-operation key — poll until terminal (async convention in references/oci-raw-request.md; track via aidp-observability).
  5. Verify with aidp schema list-tables / GET /tables?…; report the fully-qualified table name and row/column summary.

Gotchas (documented limits, no workaround)

  • Delimited files: comma only — auto-populate "Doesn't support delimiters other than comma" (platform reference §42 Known Issues #15). Pre-convert tab/pipe/semicolon-delimited files to CSV before ingest.
  • No multi-line JSON for external tables — "Can't create external tables with multi-line JSON" (platform reference §42 Known Issues #12). Use newline-delimited JSON (one record per line) for external tables.

Notes

  • Big files: prefer landing into a volume / object storage and loading from there; mind cluster memory.
  • For continuous/streaming or external-source ingestion, use the spark-connectors plugin + aidp-federate, not this skill (this is file→table).
  • Clean up temporary tables created during validation.

References

  • references/aidp-cli-map.md · references/payloads.md · references/oci-raw-request.md · references/rest-endpoint-map.md
  • pairs with aidp-workspace-files, aidp-volumes, aidp-profiling-tables

Source and attribution

Source:oracle-samples/oracle-aidp-samplesinai/claude-code-plugins/oracle-ai-data-platform-workbench-engineer-agent/skills/aidp-ingest-file-to-tableat commit90b42d6

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from oracle-samples/oracle-aidp-samples

Aidp Workspace Admin

oracle-samples

Provision and inspect AIDP DataLake instances and workspaces, including private-network workspaces attached to a customer VCN/subnet. Use when the user wants to create/list/get a workspace or DataLake instance, set up a new (e.g. private) AIDP environment, or replicate a customer setup. Create/delete are guarded — confirm before any provisioning.

Awaiting classificationOct 8, 2026

Aidp Volumes

oracle-samples

Work with AIDP volumes — list volumes, browse files inside a volume, upload/download via the PAR flow, and create directories. Use when the user mentions volumes, needs to stage large/binary files, or move data in/out of a volume (distinct from the workspace filesystem). Control-plane via the official `aidp` CLI.

Awaiting classificationOct 8, 2026

Aidp Verified Queries

oracle-samples

Maintains a repository of validated question-to-Spark-SQL pairs so an agent reuses trusted SQL before writing new queries.

Data & AnalyticsOct 8, 2026

Aidp User Settings

oracle-samples

Manage AIDP DataLake user settings and preferences via the aidp CLI or oci raw-request fallback.

Productivity & WorkflowOct 8, 2026

Aidp Spark Optimization

oracle-samples

Guides Apache Spark 3.5.0 performance tuning: partitions, shuffle, joins, skew, memory, file layout, AQE and Delta Lake.

Data & AnalyticsOct 8, 2026

Aidp Semantic Model

oracle-samples

Maintains a .aidp/semantic.md business-meaning layer defining metrics, joins, synonyms and value dictionaries for NL-to-SQL grounding.

Data & AnalyticsOct 8, 2026