Data Quality Frameworks

作者 wshobson46891e7e60da無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Implement data quality validation with Great Expectations, dbt tests, and data contracts. Use when building data quality pipelines, implementing validation rules, or establishing data contracts.

僅含說明Data & Analytics
AI 產生的概覽

指導使用 Great Expectations、dbt 測試與資料契約來實施資料品質驗證。

功能
此技能提供使用 Great Expectations、dbt 測試與資料契約來驗證資料管線的正式環境模式。內容涵蓋完整性、唯一性、有效性、準確性、一致性與時效性等資料品質構面,以及資料測試金字塔。它包含安裝指令、期望套件範例、一個跨資料表執行驗證並產生通過/失敗報告的管線類別,以及最佳實務建議。更詳細的模式文件會以另一個檔案提供。
適用情境
適用於在管線中加入資料品質檢查、設定 Great Expectations 驗證、建立 dbt 測試套件,或在團隊之間建立資料契約。也適合監控資料品質指標,以及在 CI/CD 中自動化驗證。
執行需求
需要 Python 與 great_expectations 套件,dbt 測試另需 dbt;此技能僅為說明文件,不隨附指令碼。它會引用另一個詳細檔案以提供更深入的模式。

Data Quality Frameworks

Production patterns for implementing data quality with Great Expectations, dbt tests, and data contracts to ensure reliable data pipelines.

When to Use This Skill

  • Implementing data quality checks in pipelines
  • Setting up Great Expectations validation
  • Building comprehensive dbt test suites
  • Establishing data contracts between teams
  • Monitoring data quality metrics
  • Automating data validation in CI/CD

Core Concepts

1. Data Quality Dimensions

DimensionDescriptionExample Check
CompletenessNo missing valuesexpect_column_values_to_not_be_null
UniquenessNo duplicatesexpect_column_values_to_be_unique
ValidityValues in expected rangeexpect_column_values_to_be_in_set
AccuracyData matches realityCross-reference validation
ConsistencyNo contradictionsexpect_column_pair_values_A_to_be_greater_than_B
TimelinessData is recentexpect_column_max_to_be_between

2. Testing Pyramid for Data

          /\         /  \     Integration Tests (cross-table)        /────\       /      \   Unit Tests (single column)      /────────\     /          \ Schema Tests (structure)    /────────────\

Quick Start

Great Expectations Setup

bash
# Installpip install great_expectations
# Initialize projectgreat_expectations init
# Create datasourcegreat_expectations datasource new
python
# great_expectations/checkpoints/daily_validation.ymlimport great_expectations as gx
# Create contextcontext = gx.get_context()
# Create expectation suitesuite = context.add_expectation_suite("orders_suite")
# Add expectationssuite.add_expectation(    gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id"))suite.add_expectation(    gx.expectations.ExpectColumnValuesToBeUnique(column="order_id"))
# Validateresults = context.run_checkpoint(checkpoint_name="daily_orders")

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Summary: {total_passed}/{total_tables} tables passed")

    report.append("")
    for table, result in results.items():        status = "✅" if result.passed else "❌"        report.append(f"### {status} {table}")        report.append(f"- Expectations: {result.total_expectations}")        report.append(f"- Failed: {result.failed_expectations}")
        if not result.passed:            report.append("- Failed checks:")            for detail in result.details:                if not detail["success"]:                    report.append(f"  - {detail['expectation']}: {detail['observed_value']}")        report.append("")
    return "\n".join(report)

Usage

context = gx.get_context() pipeline = DataQualityPipeline(context)

tables_to_validate = { "orders": "orders_suite", "customers": "customers_suite", "products": "products_suite", }

results = pipeline.run_all(tables_to_validate) report = pipeline.generate_report(results)

Fail pipeline if any table failed

if not all(r.passed for r in results.values()): print(report) raise ValueError("Data quality checks failed!")


## Best Practices
### Do's
- **Test early** - Validate source data before transformations- **Test incrementally** - Add tests as you find issues- **Document expectations** - Clear descriptions for each test- **Alert on failures** - Integrate with monitoring- **Version contracts** - Track schema changes
### Don'ts
- **Don't test everything** - Focus on critical columns- **Don't ignore warnings** - They often precede failures- **Don't skip freshness** - Stale data is bad data- **Don't hardcode thresholds** - Use dynamic baselines- **Don't test in isolation** - Test relationships too

來源與署名

來源:wshobson/agents位於plugins/data-engineering/skills/data-quality-frameworks提交46891e7

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架