
Aidp Spark Optimization
by oracle-samples90b42d6c24d4No licenseListed Oct 8, 2026Updated Oct 8, 2026
Use when a Spark job/notebook is slow, missing an SLA, spilling, OOMing, generating too many small files, or shuffling/skewing heavily; when reviewing Spark/PySpark code or a Spark UI for performance; or before running a large Spark workload. Covers open-source Apache Spark 3.5.0 (+ Delta Lake) tuning -- partitions, shuffle, joins, skew, file layout, memory, codegen, caching, AQE, compression, and configuration.
Add to a SourceWeft workspace
- Open the skill in your dashboard and add it to a workspace.
- Enable it for the chats that should use it.
This skill is instructions only: it ships no scripts to execute.
Add to SourceWeftYou will be asked to sign in first, then taken straight to this skill.
Ask your agent to install it
Paste this prompt into Claude Code, Codex, Cursor or another agent that can run commands — or into SourceWeft chat. The agent reads this skill's install guide, shows you its source, license and scripts, and installs it with the SourceWeft CLI once you agree.
Read https://sourceweft.com/skills/gh-oracle-samples-oracle-aidp-samples-aidp-spark-optimization/install.md and install the skill it describes. Before installing, show me its source, license and whether it ships scripts, and wait for my OK. Ask me before changing anything else on my machine.Install it yourself from a terminal
For Claude Code, Codex, Cursor and other local agents. The SourceWeft CLI fetches the skill from its source repository at the commit scanned here, and verifies every file against the hashes recorded when the skill was scanned. If anything differs, nothing is written.
npx @sourceweft/cli skills install gh-oracle-samples-oracle-aidp-samples-aidp-spark-optimizationAdd --agent claude-code, codex, cursor or universal to choose which agent gets it (Claude Code by default).
Upstream installer — not verified by SourceWeft
The open-source skills installer fetches the same pinned commit, but does not check the files against the hashes SourceWeft recorded.
npx skills add https://github.com/oracle-samples/oracle-aidp-samples/tree/90b42d6c24d4e5da784842c21234a750f7e55fef/ai/claude-code-plugins/oracle-ai-data-platform-workbench-engineer-agent/skills/aidp-spark-optimizationSource and attribution
Source:oracle-samples/oracle-aidp-samplesinai/claude-code-plugins/oracle-ai-data-platform-workbench-engineer-agent/skills/aidp-spark-optimizationat commit90b42d6
License: No license
Content belongs to its original authors. SourceWeft indexes it from a public repository.
More from oracle-samples/oracle-aidp-samples

Aidp Workspace Admin
oracle-samples
Provision and inspect AIDP DataLake instances and workspaces, including private-network workspaces attached to a customer VCN/subnet. Use when the user wants to create/list/get a workspace or DataLake instance, set up a new (e.g. private) AIDP environment, or replicate a customer setup. Create/delete are guarded — confirm before any provisioning.

Aidp Volumes
oracle-samples
Work with AIDP volumes — list volumes, browse files inside a volume, upload/download via the PAR flow, and create directories. Use when the user mentions volumes, needs to stage large/binary files, or move data in/out of a volume (distinct from the workspace filesystem). Control-plane via the official `aidp` CLI.

Aidp Verified Queries
oracle-samples
Maintains a repository of validated question-to-Spark-SQL pairs so an agent reuses trusted SQL before writing new queries.

Aidp User Settings
oracle-samples
Manage AIDP DataLake user settings and preferences via the aidp CLI or oci raw-request fallback.

Aidp Semantic Model
oracle-samples
Maintains a .aidp/semantic.md business-meaning layer defining metrics, joins, synonyms and value dictionaries for NL-to-SQL grounding.

Aidp Roles Access
oracle-samples
Manages AIDP roles and access: list roles, add or remove members, and grant or revoke per-resource permissions.
More in Data & Analytics

Spanner Basics
Guides Google Cloud Spanner administration, schema design, querying and performance diagnosis.

Gke Cost Analysis
Answers natural-language questions about GKE cluster and workload costs using BigQuery billing exports and live cluster metrics.

Datalineage Bigquery Asset Impact Analysis
Guides an agent through downstream impact (blast radius) analysis for a BigQuery table or view using Data Lineage.

Bigquery Troubleshooting
Diagnoses failing, slow, or unexpectedly expensive BigQuery jobs through structured root-cause workflows.

Bigquery Optimization
Guides BigQuery cost and performance optimization across capacity editions, storage layout, and SQL queries.

Bigquery Bigframes
Guides writing Python code with BigQuery DataFrames (BigFrames) for data processing, analysis, and machine learning.