Profiling Tables

作者 astronomercbe1141f547b無授權條款451 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Deep-dive data profiling for a specific table. Use when the user asks to profile a table, wants statistics about a dataset, asks about data quality, or needs to understand a table's structure and content. Requires a table name.

僅含說明Data & Analytics
AI 產生的概覽

透過 SQL 統計、基數分析、樣本資料與資料品質評估,對資料庫資料表進行剖析。

功能
引導代理對指定的單一資料表進行結構化 SQL 剖析:欄位中繼資料、資料列數、依型別計算的欄位統計、基數與高頻值,以及樣本資料列。接著從完整性、唯一性、新鮮度、有效性與一致性進行評估。最終產出包含概觀、欄位結構表、關鍵統計、資料品質評分與建議查詢的書面剖析報告。
適用情境
當使用者要求剖析某個資料表、想了解資料集的統計資訊或資料品質,或需要理解陌生資料表的結構與內容時使用。必須提供資料表名稱。
執行需求
需要可存取的 SQL 資料庫,以及用於執行查詢的 run_sql 工具;資料表名稱需明確,或可透過 INFORMATION_SCHEMA 定位。此技能不附帶指令碼。

Data Profile

Generate a comprehensive profile of a table that a new team member could use to understand the data.

Step 1: Basic Metadata

Query column metadata:

sql
SELECT COLUMN_NAME, DATA_TYPE, COMMENTFROM <database>.INFORMATION_SCHEMA.COLUMNSWHERE TABLE_SCHEMA = '<schema>' AND TABLE_NAME = '<table>'ORDER BY ORDINAL_POSITION

If the table name isn't fully qualified, search INFORMATION_SCHEMA.TABLES to locate it first.

Step 2: Size and Shape

Run via run_sql:

sql
SELECT    COUNT(*) as total_rows,    COUNT(*) / 1000000.0 as millions_of_rowsFROM <table>

Step 3: Column-Level Statistics

For each column, gather appropriate statistics based on data type:

Numeric Columns

sql
SELECT    MIN(column_name) as min_val,    MAX(column_name) as max_val,    AVG(column_name) as avg_val,    STDDEV(column_name) as std_dev,    PERCENTILE_CONT(0.5) WITHIN GROUP (ORDER BY column_name) as median,    SUM(CASE WHEN column_name IS NULL THEN 1 ELSE 0 END) as null_count,    COUNT(DISTINCT column_name) as distinct_countFROM <table>

String Columns

sql
SELECT    MIN(LEN(column_name)) as min_length,    MAX(LEN(column_name)) as max_length,    AVG(LEN(column_name)) as avg_length,    SUM(CASE WHEN column_name IS NULL OR column_name = '' THEN 1 ELSE 0 END) as empty_count,    COUNT(DISTINCT column_name) as distinct_countFROM <table>

Date/Timestamp Columns

sql
SELECT    MIN(column_name) as earliest,    MAX(column_name) as latest,    DATEDIFF('day', MIN(column_name), MAX(column_name)) as date_range_days,    SUM(CASE WHEN column_name IS NULL THEN 1 ELSE 0 END) as null_countFROM <table>

Step 4: Cardinality Analysis

For columns that look like categorical/dimension keys:

sql
SELECT    column_name,    COUNT(*) as frequency,    ROUND(COUNT(*) * 100.0 / SUM(COUNT(*)) OVER(), 2) as percentageFROM <table>GROUP BY column_nameORDER BY frequency DESCLIMIT 20

This reveals:

  • High-cardinality columns (likely IDs or unique values)
  • Low-cardinality columns (likely categories or status fields)
  • Skewed distributions (one value dominates)

Step 5: Sample Data

Get representative rows:

sql
SELECT *FROM <table>LIMIT 10

If the table is large and you want variety, sample from different time periods or categories.

Step 6: Data Quality Assessment

Summarize quality across dimensions:

Completeness

  • Which columns have NULLs? What percentage?
  • Are NULLs expected or problematic?

Uniqueness

  • Does the apparent primary key have duplicates?
  • Are there unexpected duplicate rows?

Freshness

  • When was data last updated? (MAX of timestamp columns)
  • Is the update frequency as expected?

Validity

  • Are there values outside expected ranges?
  • Are there invalid formats (dates, emails, etc.)?
  • Are there orphaned foreign keys?

Consistency

  • Do related columns make sense together?
  • Are there logical contradictions?

Step 7: Output Summary

Provide a structured profile:

Overview

2-3 sentences describing what this table contains, who uses it, and how fresh it is.

Schema

ColumnTypeNulls%DistinctDescription
...............

Key Statistics

  • Row count: X
  • Date range: Y to Z
  • Last updated: timestamp

Data Quality Score

  • Completeness: X/10
  • Uniqueness: X/10
  • Freshness: X/10
  • Overall: X/10

Potential Issues

List any data quality concerns discovered.

Recommended Queries

3-5 useful queries for common questions about this data.

來源與署名

來源:astronomer/agents位於skills/profiling-tables提交cbe1141

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架