Aidp Data Quality

by oracle-samples90b42d6c24d4No licenseListed Oct 8, 2026Updated Oct 8, 2026

Run data-quality rule checks on AIDP tables — not-null, uniqueness, allowed ranges/sets, referential integrity, and freshness. Use when the user wants to validate data, check for nulls/duplicates/orphans, assert a column's domain, or gate a pipeline on quality. Expresses each rule as bounded Spark SQL and reports pass/fail with offending counts.

Instructions onlyData & Analytics
AI-generated overview

Validates AIDP tables against data-quality rules expressed as bounded Spark SQL, reporting pass/fail and violation counts.

What it does
Compiles explicit data-quality rules — not-null, uniqueness, allowed ranges or sets, referential integrity, and freshness — into bounded Spark SQL violation-count queries. Runs each query against AIDP tables through a bundled helper script and treats a count of zero as PASS. On failures it samples a few offending rows and reports a summary table of rule, target, result, and violation count. It can also persist validated rules for re-runs and wire checks into a pipeline as a gating task.
When to use it
Use when someone wants to validate a table or column, check for nulls, duplicates, orphan rows, or stale data, assert a column's allowed domain, or gate a pipeline on data quality.
Requirements
Requires access to an AIDP environment: a running cluster, a region, a datalake OCID, a workspace, and a cluster key, plus an api_key DEFAULT profile for minting a UPST. It relies on a bundled helper script (scripts/aidp_sql.py) and reference documents, and optionally on .aidp/catalog.md, .aidp/semantic.md, and .aidp/dq-rules.md. Network access to the AIDP service is needed. The skill itself ships no scripts of its own.

aidp-data-quality — rule checks via Spark SQL

Validate AIDP tables against explicit data-quality rules, each compiled to bounded Spark SQL and executed with the bundled helper — no MCP and no ai-data-engineer-agent repo required.

When to use

  • "Check <table> for nulls/duplicates", "validate <column>", "are there orphan rows", "is the data fresh", or gating a pipeline on quality.

Rule types (each → a counting SQL that should return 0 violations)

RuleCheck (violations)
not-nullCOUNT(*) WHERE col IS NULL
uniqueCOUNT(*) - COUNT(DISTINCT key) (or GROUP BY key HAVING COUNT(*)>1)
range / setCOUNT(*) WHERE col NOT BETWEEN lo AND hi / col NOT IN (...)
referentialCOUNT(*) child LEFT JOIN parent ... WHERE parent.key IS NULL
freshnessMAX(ts) vs SLA (e.g. datediff(current_date, MAX(ts)) <= N)

Workflow

  1. Resolve table(s)/columns; use join keys from .aidp/catalog.md for referential checks (don't guess). Pull rule definitions from .aidp/semantic.md value dictionaries where available.
  2. Ensure the cluster is RUNNING (aidp-cluster-ops / oci raw-request), then for each rule run the violation-count SQL with the bundled helper (PASS if 0, else FAIL):
    bash
    python "$PLUGIN_DIR/scripts/aidp_sql.py" --region <region> --datalake <DATALAKE_OCID> --workspace <ws> \  --cluster <cluster-key> \  --code "spark.sql('''SELECT COUNT(*) AS v FROM cat.sch.t WHERE col IS NULL''').show()"
    It mints a UPST from the api_key DEFAULT profile, auto-creates a scratch notebook, and returns JSON with status / outputs / spark_job_ids. No AIDP_SESSION required (--session-profile optional).
  3. On a non-zero count, FAIL and pull a few example offending rows with a separate bounded LIMIT query.
  4. Report a summary table: rule · target · result · violation count.
  5. Offer to (a) persist the rule set for re-runs (see below), and (b) wire checks into a Job (aidp-pipelines) as a gating task.

Persisting a re-runnable rule set

Register validated rules in .aidp/dq-rules.md so they can be re-run later (the quality analogue of .aidp/verified-queries.md). One entry per rule records the target table/column, rule-type (the five types above), the violation-SQL (counts violations → PASS when 0), and last-result / last-checked. To re-run, execute each entry's stored violation-SQL via scripts/aidp_sql.py, set the result to PASS (0) or FAIL (<count>), and record the cluster + date — never mark PASS without a status: ok run returning 0. Format and re-run rules: references/dq-rules.md.

Reliability rules

  • Run real SQL via scripts/aidp_sql.py; never assert a rule passed without a status: ok result.
  • Keep checks bounded; sample example offenders rather than dumping full result sets.
  • If a cell returns status: error, read the error, fix the SQL grounded in the catalog, and retry.

References

  • references/dq-rules.md (.aidp/dq-rules.md rule-set format + re-run)
  • scripts/aidp_sql.py · references/no-mcp-rest-map.md · references/oci-raw-request.md · references/semantic-model.md

Source and attribution

Source:oracle-samples/oracle-aidp-samplesinai/claude-code-plugins/oracle-ai-data-platform-workbench-engineer-agent/skills/aidp-data-qualityat commit90b42d6

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from oracle-samples/oracle-aidp-samples

Aidp Workspace Admin

oracle-samples

Provision and inspect AIDP DataLake instances and workspaces, including private-network workspaces attached to a customer VCN/subnet. Use when the user wants to create/list/get a workspace or DataLake instance, set up a new (e.g. private) AIDP environment, or replicate a customer setup. Create/delete are guarded — confirm before any provisioning.

Awaiting classificationOct 8, 2026

Aidp Volumes

oracle-samples

Work with AIDP volumes — list volumes, browse files inside a volume, upload/download via the PAR flow, and create directories. Use when the user mentions volumes, needs to stage large/binary files, or move data in/out of a volume (distinct from the workspace filesystem). Control-plane via the official `aidp` CLI.

Awaiting classificationOct 8, 2026

Aidp Verified Queries

oracle-samples

Maintains a repository of validated question-to-Spark-SQL pairs so an agent reuses trusted SQL before writing new queries.

Data & AnalyticsOct 8, 2026

Aidp User Settings

oracle-samples

Manage AIDP DataLake user settings and preferences via the aidp CLI or oci raw-request fallback.

Productivity & WorkflowOct 8, 2026

Aidp Spark Optimization

oracle-samples

Guides Apache Spark 3.5.0 performance tuning: partitions, shuffle, joins, skew, memory, file layout, AQE and Delta Lake.

Data & AnalyticsOct 8, 2026

Aidp Semantic Model

oracle-samples

Maintains a .aidp/semantic.md business-meaning layer defining metrics, joins, synonyms and value dictionaries for NL-to-SQL grounding.

Data & AnalyticsOct 8, 2026