Aidp Cluster Ops

作者 oracle-samples90b42d6c24d4无许可证收录于 2026年10月8日更新于 2026年10月8日

Manage AIDP Spark compute clusters — list, status, start/stop/restart, installed libraries (JARs/Python), provision/scale a new cluster (driver/worker shapes, autoscale, GPU/RAPIDS, AI Compute), and connect external BI tools (JDBC/ODBC). Use when the user asks about clusters, needs to start/stop compute, create or scale a cluster, install libraries, set up a GPU cluster, use AI Compute for agent flows, connect Tableau/Power BI/DBeaver, or pick a cluster before running data work.

仅含说明DevOps & Cloud
AI 生成的概览

管理 AIDP Spark 集群:列表、状态、启动/停止/重启、库、创建与扩缩容以及 BI 连接。

功能
该控制平面技能通过官方 aidp CLI 检查和控制 AIDP Spark 计算集群,并以 oci raw-request REST 调用作为回退方案。它涵盖列出集群、查看状态与就绪情况、启动、停止和重启、列出或修补已安装的 JAR 与 Python 库,以及使用驱动与工作节点规格、自动扩缩容和空闲超时来创建或扩缩容集群。它还说明了 GPU/RAPIDS 配置、用于代理流程的 AI Compute,以及外部 BI 工具的 JDBC/ODBC 连接信息。
适用场景
当用户询问有哪些集群、计算是否在运行,或需要在数据工作前启动、停止或重启集群时使用。它也适用于创建或扩缩容集群、安装库、搭建 GPU 集群,或连接 Tableau、Power BI、DBeaver。
运行要求
需要以 OCI 身份验证访问 AIDP 数据平面 REST API,可使用官方 aidp CLI 或用于原始请求的 oci CLI,并提供实例 ID、区域和配置文件。需要能访问 AIDP 端点的网络。该技能不附带脚本,仅为说明文档。

aidp-cluster-ops — cluster lifecycle & libraries

Inspect and control AIDP Spark clusters. Most data skills depend on a RUNNING cluster, so this is the common pre-step. This is a control-plane skill. No MCP and no ai-data-engineer-agent repo are required.

When to use

  • "What clusters are there / is X running?", "start/stop/restart the cluster", "what libraries are installed", "check if compute is up before running data work".

Engine — official aidp CLI (control-plane)

Preferred engine is the official Oracle aidp CLI; oci raw-request is the fallback when the CLI isn't installed. Both hit the same data-plane REST API with the same auth — see references/aidp-cli-map.md for the full command map, references/oci-raw-request.md for base URL + auth ladder + async/error conventions, and references/no-mcp-rest-map.md for REST endpoint shapes.

OpCLI (preferred)REST fallback
List clustersaidp cluster listGET /clusters — or GET /workspaces/<ws>/clusters
Status / configaidp cluster get --cluster-key <key>GET /workspaces/<ws>/clusters/<key>
Default clusteraidp cluster get-default(in list output)
Librariesaidp cluster list-libraries --cluster-key <key> (patch-library to add/remove)inside cluster GET
Start / Stop / Restartaidp cluster start|stop|restart --cluster-key <key>POST /workspaces/<ws>/clusters/<key>/actions/start|stop|restart
Logs / metricsaidp cluster search-logs|download-logs|summarize-metrics-data --cluster-key <key>…/clusters/<key>/…

All CLI calls take --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region <r>.

bash
# CLI (preferred): list clustersaidp cluster list --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region us-ashburn-1
# CLI (preferred): cluster detail — state, config, connectionsaidp cluster get --cluster-key <KEY> --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region us-ashburn-1
# CLI (preferred): start (the CLI handles the required body for you)aidp cluster start --cluster-key <KEY> --instance-id <DATALAKE_OCID> --auth api_key --profile DEFAULT --region us-ashburn-1

Mutating ops (start/stop/restart, patch-library) — for shared clusters, persist any non-trivial body to .aidp/payloads/ and confirm first (references/payloads.md).

Fallback (no CLI installed) — oci raw-request against https://aidp.<region>.oci.oraclecloud.com/20240831/dataLakes/<DATALAKE_OCID>/…:

bash
# List clusters (verified GET)oci raw-request --http-method GET \  --target-uri "https://aidp.us-ashburn-1.oci.oraclecloud.com/20240831/dataLakes/<OCID>/workspaces/<WS>/clusters" \  --profile DEFAULT
# Cluster detail — state, config, connections, AND installed librariesoci raw-request --http-method GET \  --target-uri "https://aidp.us-ashburn-1.oci.oraclecloud.com/20240831/dataLakes/<OCID>/workspaces/<WS>/clusters/<KEY>" \  --profile DEFAULT
# Start (POST action) — a JSON body is REQUIRED (use {}); empty body 400soci raw-request --http-method POST \  --target-uri "https://aidp.us-ashburn-1.oci.oraclecloud.com/20240831/dataLakes/<OCID>/workspaces/<WS>/clusters/<KEY>/actions/start" \  --request-body '{}' --request-headers '{"content-type":"application/json"}' \  --profile DEFAULT

Patterns

  • Find compute: aidp cluster list (REST GET /clusters) lists DataLake clusters; GET /workspaces/<ws>/clusters scopes the REST fallback to one workspace. Cross-workspace questions must pass the right <ws> — don't rely on a single default workspace.
  • Status + readiness: aidp cluster get --cluster-key <key> (REST GET /workspaces/<ws>/clusters/<key>) returns state (e.g. STARTING → ACTIVE) + stateDetails, config, connections, and attached notebooks/sessions. Data/SQL skills should call this first and start if stopped, then re-check, polling state until ACTIVE (start takes minutes).
  • Lifecycle: aidp cluster start|stop|restart (REST …/actions/start|stop|restart). The async 202 is poll-to-terminal. Confirm before stopping a shared cluster.
  • Libraries: aidp cluster list-libraries --cluster-key <key>; with the REST fallback the installed JARs + Python libs come back inside the cluster GET. Check before relying on a connector/lib.

Caveats

  • REST fallback only — actions/start|stop|restart need a body {} (LIVE-VERIFIED 2026-06-09): calling with no body returns 400 InvalidParameter: The request body must not be null; passing --request-body '{}' returns 202. This (not workspace mismatch) was the original "start 400". The aidp CLI sets this body for you. A second start while already STARTING returns 409 Conflict (expected) on either engine.
  • Use the cluster's home workspace in the REST action URL — find it via GET /workspaces/<ws>/clusters (a cluster may not live in your default workspace). The CLI resolves the workspace from the cluster key.
  • Per the no-fabrication gate in oci-raw-request.md: don't present an endpoint/version/prefix as confirmed until a live 2xx (or documented 4xx) is recorded in rest-endpoint-map.md.

Provision / scale a cluster (aidp cluster create|update|delete; REST POST/PUT/DELETE …/workspaces/<ws>/clusters)

Live-verified create body (this provisioned agent_e2e_cluster → ACTIVE, 2026-06-10):

json
{ "type": "USER", "displayName": "etl_cluster",  "driverConfig": { "driverShape": "amd.generic", "driverShapeConfig": { "ocpus": 2, "memoryInGBs": 16, "gpus": 0 } },  "workerConfig": { "minWorkerCount": 1, "maxWorkerCount": 1, "workerShape": "amd.generic",                    "workerShapeConfig": { "ocpus": 2, "memoryInGBs": 16, "gpus": 0 } },  "clusterRuntimeConfig": { "sparkVersion": "3.5.0", "type": "SPARK", "initScripts": [] },  "autoTerminationMinutes": 120 }
  • displayName charset: must start with a letter; the only special chars allowed are underscore and slash. A hyphen (e.g. etl-cluster) → 400 InvalidParameter ("no special characters … except for underscore, slash") — use etl_cluster.
  • Shapes: AMD / ARM / Intel / NVIDIA GPU (platform-ref §12). Quickstart = 1 driver + ≤10 workers, AMD 2 OCPU/32 GB, autoscale (fast start); Custom = full control. Installing custom libs to a Quickstart cluster converts it to Custom.
  • Scale: static (minWorkerCount == maxWorkerCount) or autoscale (min < max).
  • Run duration: always-on (omit) or idle timeout via autoTerminationMinutes.
  • Runtime: Spark 3.5.0 / Delta 3.2.0 / Python 3.11 / Java 17 (Python + SQL user code only).

Libraries (install) — aidp cluster patch-library

Formats .jar / .whl / requirements.txt; source = workspace / volume / uploaded file. Must restart the cluster after installing. Notebook-scoped installs (!pip install …, .ipynb only) don't need a restart but apply only to that notebook (see aidp-notebooks).

GPU / RAPIDS clusters (platform-ref §14)

GPU shapes: 1 GPU = 15 OCPU/24 GB GPU mem; 2 GPU = 30 OCPU/48 GB. Rule: both driver AND worker must be NVIDIA GPU — no CPU/GPU mixing. Required RAPIDS Spark configs: spark.plugins=com.nvidia.spark.SQLPlugin, spark.shuffle.manager=com.nvidia.spark.rapids.spark350.RapidsShuffleManager, spark.rapids.shuffle.mode=MULTITHREADED, spark.executor.resource.gpu.amount=1, spark.task.resource.gpu.amount=1/executor.cores. Libraries: Spark RAPIDS, Spark RAPIDS ML (cuML).

AI Compute (Preview) — powers agent flows

Specialized compute for agent flows (aidp-agent-flows / aidp-agent-highcode): 1–64 OCPU, restricted to PvtDefaultWorkspace; private workspaces must connect to a private Autonomous AI Lakehouse (immutable once linked); public workspaces can't use private ALH (AI features unavailable). Create via Workspace > Create > AI Compute; Start/Stop frees/meters compute; attached flows show on the cluster's Agent flows tab.

Connect external BI tools (JDBC / ODBC) — additive to OAC/FDI, not a replacement

The cluster Connection Details tab provides the simbaSpark JDBC and ODBC drivers for DBeaver / Tableau / Power BI. JDBC driver class com.simba.spark.jdbc.Driver; use the JDBC URL from that tab. Auth: token-based (no ociProfile in the URL → browser SSO) or API key (append ociProfile=<profile_name>). OAC connection setup itself is OAC-side (the OAC/fusion-bundle plugin); here we only expose the AIDP driver/URL.

Notes

  • For Spark job/stage/task diagnostics on a running cluster, use aidp-spark-debugging.
  • Optional accelerator: if an aidp MCP happens to be configured, its list_clusters / get_cluster_status / get_default_cluster / start_cluster / stop_cluster / restart_cluster / list_cluster_libraries tools wrap these same REST calls. The MCP is not required — the oci raw-request calls above are the source of truth.

References

  • references/aidp-cli-map.md — skill → official aidp CLI command map (primary engine)
  • references/oci-raw-request.md — base URL, auth ladder, async/errors
  • references/no-mcp-rest-map.md — cluster endpoint map + start-400 note
  • references/rest-endpoint-map.md — verification ledger

来源与署名

来源:oracle-samples/oracle-aidp-samples位于ai/claude-code-plugins/oracle-ai-data-platform-workbench-engineer-agent/skills/aidp-cluster-ops提交90b42d6

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 oracle-samples/oracle-aidp-samples 的技能

Aidp Workspace Admin

oracle-samples

Provision and inspect AIDP DataLake instances and workspaces, including private-network workspaces attached to a customer VCN/subnet. Use when the user wants to create/list/get a workspace or DataLake instance, set up a new (e.g. private) AIDP environment, or replicate a customer setup. Create/delete are guarded — confirm before any provisioning.

待分类2026年10月8日

Aidp Volumes

oracle-samples

Work with AIDP volumes — list volumes, browse files inside a volume, upload/download via the PAR flow, and create directories. Use when the user mentions volumes, needs to stage large/binary files, or move data in/out of a volume (distinct from the workspace filesystem). Control-plane via the official `aidp` CLI.

待分类2026年10月8日

Aidp Verified Queries

oracle-samples

维护经过验证的问题到 Spark SQL 配对库,让智能体在生成新 SQL 前优先复用可信查询。

Data & Analytics2026年10月8日

Aidp User Settings

oracle-samples

通过 aidp CLI 或 oci raw-request 备用方式管理 AIDP DataLake 用户设置与偏好。

Productivity & Workflow2026年10月8日

Aidp Spark Optimization

oracle-samples

指导 Apache Spark 3.5.0 性能调优:分区、shuffle、连接、倾斜、内存、文件布局、AQE 与 Delta Lake。

Data & Analytics2026年10月8日

Aidp Semantic Model

oracle-samples

维护 .aidp/semantic.md 业务语义层,定义指标、连接、同义词和值字典,为自然语言转 SQL 提供依据。

Data & Analytics2026年10月8日