aidp-data-quality — rule checks via Spark SQL
Validate AIDP tables against explicit data-quality rules, each compiled to bounded Spark SQL and executed
with the bundled helper — no MCP and no ai-data-engineer-agent repo required.
When to use
- "Check <table> for nulls/duplicates", "validate <column>", "are there orphan rows", "is the data fresh", or gating a pipeline on quality.
Rule types (each → a counting SQL that should return 0 violations)
Workflow
- Resolve table(s)/columns; use join keys from
.aidp/catalog.mdfor referential checks (don't guess). Pull rule definitions from.aidp/semantic.mdvalue dictionaries where available. - Ensure the cluster is RUNNING (
aidp-cluster-ops/oci raw-request), then for each rule run the violation-count SQL with the bundled helper (PASS if 0, else FAIL): It mints a UPST from the api_key DEFAULT profile, auto-creates a scratch notebook, and returns JSON withstatus/outputs/spark_job_ids. NoAIDP_SESSIONrequired (--session-profileoptional). - On a non-zero count, FAIL and pull a few example offending rows with a separate bounded
LIMITquery. - Report a summary table: rule · target · result · violation count.
- Offer to (a) persist the rule set for re-runs (see below), and (b) wire checks into a Job
(
aidp-pipelines) as a gating task.
Persisting a re-runnable rule set
Register validated rules in .aidp/dq-rules.md so they can be re-run later (the quality analogue of
.aidp/verified-queries.md). One entry per rule records the target table/column, rule-type (the five types
above), the violation-SQL (counts violations → PASS when 0), and last-result / last-checked. To
re-run, execute each entry's stored violation-SQL via scripts/aidp_sql.py, set the result to PASS (0) or
FAIL (<count>), and record the cluster + date — never mark PASS without a status: ok run returning 0.
Format and re-run rules: references/dq-rules.md.
Reliability rules
- Run real SQL via
scripts/aidp_sql.py; never assert a rule passed without astatus: okresult. - Keep checks bounded; sample example offenders rather than dumping full result sets.
- If a cell returns
status: error, read the error, fix the SQL grounded in the catalog, and retry.
References
- references/dq-rules.md (
.aidp/dq-rules.mdrule-set format + re-run) - scripts/aidp_sql.py · references/no-mcp-rest-map.md · references/oci-raw-request.md · references/semantic-model.md


