All skills
wshobson avatar

/data-quality-frameworks

@be57c0b
by Seth Hobsonwshobson/agents40k stars
4,281

Implement data quality validation with Great Expectations, dbt tests, and data contracts. Use when building data quality pipelines, implementing validation rules, or establishing data contracts.

Use this Skill: https://skilld.dev/gh/wshobson/agents/data-quality-frameworks

This session only. Nothing lands on disk.

SKILL.md

โ‰ˆ55 tokens always: the name and description. โ‰ˆ1.1k when used: this file. โ‰ˆ2.9k more on demand in 1 file.

Data Quality Frameworks

Production patterns for implementing data quality with Great Expectations, dbt tests, and data contracts to ensure reliable data pipelines.

When to Use This Skill

  • Implementing data quality checks in pipelines
  • Setting up Great Expectations validation
  • Building comprehensive dbt test suites
  • Establishing data contracts between teams
  • Monitoring data quality metrics
  • Automating data validation in CI/CD

Core Concepts

1. Data Quality Dimensions

Dimension Description Example Check
Completeness No missing values expect_column_values_to_not_be_null
Uniqueness No duplicates expect_column_values_to_be_unique
Validity Values in expected range expect_column_values_to_be_in_set
Accuracy Data matches reality Cross-reference validation
Consistency No contradictions expect_column_pair_values_A_to_be_greater_than_B
Timeliness Data is recent expect_column_max_to_be_between

2. Testing Pyramid for Data

          /\
         /  \     Integration Tests (cross-table)
        /โ”€โ”€โ”€โ”€\
       /      \   Unit Tests (single column)
      /โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€\
     /          \ Schema Tests (structure)
    /โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€\

Quick Start

Great Expectations Setup

# Install
pip install great_expectations

# Initialize project
great_expectations init

# Create datasource
great_expectations datasource new
# great_expectations/checkpoints/daily_validation.yml
import great_expectations as gx

# Create context
context = gx.get_context()

# Create expectation suite
suite = context.add_expectation_suite("orders_suite")

# Add expectations
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
)

# Validate
results = context.run_checkpoint(checkpoint_name="daily_orders")

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Summary: {total_passed}/{total_tables} tables passed")

    report.append("")

    for table, result in results.items():
        status = "โœ…" if result.passed else "โŒ"
        report.append(f"### {status} {table}")
        report.append(f"- Expectations: {result.total_expectations}")
        report.append(f"- Failed: {result.failed_expectations}")

        if not result.passed:
            report.append("- Failed checks:")
            for detail in result.details:
                if not detail["success"]:
                    report.append(f"  - {detail['expectation']}: {detail['observed_value']}")
        report.append("")

    return "\n".join(report)

Usage

context = gx.get_context() pipeline = DataQualityPipeline(context)

tables_to_validate = { "orders": "orders_suite", "customers": "customers_suite", "products": "products_suite", }

results = pipeline.run_all(tables_to_validate) report = pipeline.generate_report(results)

Fail pipeline if any table failed

if not all(r.passed for r in results.values()): print(report) raise ValueError("Data quality checks failed!")


## Best Practices

### Do's

- **Test early** - Validate source data before transformations
- **Test incrementally** - Add tests as you find issues
- **Document expectations** - Clear descriptions for each test
- **Alert on failures** - Integrate with monitoring
- **Version contracts** - Track schema changes

### Don'ts

- **Don't test everything** - Focus on critical columns
- **Don't ignore warnings** - They often precede failures
- **Don't skip freshness** - Stale data is bad data
- **Don't hardcode thresholds** - Use dynamic baselines
- **Don't test in isolation** - Test relationships too

Source: SKILL.md on GitHub

1 warning16d5 checks ยท Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides standard documentation and code examples for implementing data quality validation using Great Expectations, dbt, and data contracts. No security risks or malicious patterns were identified.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW ยท No issues

  • Runlayer6mo

    1/1 file flagged

  • ZeroLeaks5mo

    Score: 93/100 ยท 2 sections analyzed

Signed by skilld at be57c0b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 4 months ago
  • Python
  • great-expectations
  • dbt
  • data-contracts
  • data-quality
  • validation
  • pipeline
  • schema

README badge

README badge for wshobson/agents/data-quality-frameworks

Implements data quality validation using Great Expectations, dbt tests, and data contracts to catch issues in production pipelines. Covers completeness, uniqueness, validity, accuracy, consistency, and timeliness checks across source data, transformations, and downstream tables.

Generated from the current SKILL.md.

Does this skill support dbt, Great Expectations, or both?
The skill covers both. It includes patterns for Great Expectations validation, dbt test suites, and data contracts, so you can use whichever framework fits your pipeline.
Can I use this for cross-table validation?
Yes. The skill includes patterns for integration tests across tables and relationship validation, not just single-column checks.
Does this skill integrate with CI/CD?
The skill mentions automating data validation in CI/CD and integrating with monitoring systems, but specific CI/CD pipeline examples are referenced in the detailed pattern documentation.
What data sources does this support?
The skill focuses on Great Expectations, which supports Postgres, Snowflake, BigQuery, and other common data warehouses, but the SKILL.md does not list specific database support.
Does this cover data contracts?
Yes. Establishing data contracts between teams is listed as a core use case, alongside Great Expectations validation and dbt tests.

Generated from the current SKILL.md. These answers refresh after source changes.