Data Quality Frameworks

作者 wshobson46891e7e60da无许可证收录于 2026年10月8日更新于 2026年10月8日

Implement data quality validation with Great Expectations, dbt tests, and data contracts. Use when building data quality pipelines, implementing validation rules, or establishing data contracts.

仅含说明Data & Analytics
AI 生成的概览

指导使用 Great Expectations、dbt 测试和数据契约实施数据质量校验。

功能
该技能提供使用 Great Expectations、dbt 测试和数据契约来校验数据管道的生产级模式。内容涵盖完整性、唯一性、有效性、准确性、一致性和及时性等数据质量维度,以及数据测试金字塔。它包含安装命令、期望套件示例、一个跨表运行校验并生成通过/失败报告的管道类,以及最佳实践建议。更详细的模式文档在单独的文件中引用。
适用场景
适用于在管道中添加数据质量检查、设置 Great Expectations 校验、构建 dbt 测试套件,或在团队之间建立数据契约。也适合监控数据质量指标以及在 CI/CD 中自动化校验。
运行要求
需要 Python 和 great_expectations 包,dbt 测试还需要 dbt;该技能仅为说明文档,不附带脚本。它引用单独的详细文件以提供更深入的模式。

Data Quality Frameworks

Production patterns for implementing data quality with Great Expectations, dbt tests, and data contracts to ensure reliable data pipelines.

When to Use This Skill

  • Implementing data quality checks in pipelines
  • Setting up Great Expectations validation
  • Building comprehensive dbt test suites
  • Establishing data contracts between teams
  • Monitoring data quality metrics
  • Automating data validation in CI/CD

Core Concepts

1. Data Quality Dimensions

DimensionDescriptionExample Check
CompletenessNo missing valuesexpect_column_values_to_not_be_null
UniquenessNo duplicatesexpect_column_values_to_be_unique
ValidityValues in expected rangeexpect_column_values_to_be_in_set
AccuracyData matches realityCross-reference validation
ConsistencyNo contradictionsexpect_column_pair_values_A_to_be_greater_than_B
TimelinessData is recentexpect_column_max_to_be_between

2. Testing Pyramid for Data

          /\         /  \     Integration Tests (cross-table)        /────\       /      \   Unit Tests (single column)      /────────\     /          \ Schema Tests (structure)    /────────────\

Quick Start

Great Expectations Setup

bash
# Installpip install great_expectations
# Initialize projectgreat_expectations init
# Create datasourcegreat_expectations datasource new
python
# great_expectations/checkpoints/daily_validation.ymlimport great_expectations as gx
# Create contextcontext = gx.get_context()
# Create expectation suitesuite = context.add_expectation_suite("orders_suite")
# Add expectationssuite.add_expectation(    gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id"))suite.add_expectation(    gx.expectations.ExpectColumnValuesToBeUnique(column="order_id"))
# Validateresults = context.run_checkpoint(checkpoint_name="daily_orders")

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Summary: {total_passed}/{total_tables} tables passed")

    report.append("")
    for table, result in results.items():        status = "✅" if result.passed else "❌"        report.append(f"### {status} {table}")        report.append(f"- Expectations: {result.total_expectations}")        report.append(f"- Failed: {result.failed_expectations}")
        if not result.passed:            report.append("- Failed checks:")            for detail in result.details:                if not detail["success"]:                    report.append(f"  - {detail['expectation']}: {detail['observed_value']}")        report.append("")
    return "\n".join(report)

Usage

context = gx.get_context() pipeline = DataQualityPipeline(context)

tables_to_validate = { "orders": "orders_suite", "customers": "customers_suite", "products": "products_suite", }

results = pipeline.run_all(tables_to_validate) report = pipeline.generate_report(results)

Fail pipeline if any table failed

if not all(r.passed for r in results.values()): print(report) raise ValueError("Data quality checks failed!")


## Best Practices
### Do's
- **Test early** - Validate source data before transformations- **Test incrementally** - Add tests as you find issues- **Document expectations** - Clear descriptions for each test- **Alert on failures** - Integrate with monitoring- **Version contracts** - Track schema changes
### Don'ts
- **Don't test everything** - Focus on critical columns- **Don't ignore warnings** - They often precede failures- **Don't skip freshness** - Stale data is bad data- **Don't hardcode thresholds** - Use dynamic baselines- **Don't test in isolation** - Test relationships too

来源与署名

来源:wshobson/agents位于plugins/data-engineering/skills/data-quality-frameworks提交46891e7

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架