Dummy Dataset

作者 phuryn8607e3b07781无许可证26K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库3周前更新

Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script). Use when creating test data, building mock datasets, or generating sample data for development and demos.

仅含说明Data & Analytics
AI 生成的概览

生成逼真的虚拟数据集,可自定义列、约束和输出格式,如 CSV、JSON、SQL 或 Python。

功能
该技能引导智能体生成合成测试数据:确定数据领域、定义列名与类型、设定行数,并应用贴近真实的模式和业务约束。它可以输出可直接运行的 Python 生成脚本,或直接输出 CSV、JSON、SQL INSERT 等数据文件,并附带校验说明和快速上手说明。产出用于开发、演示和测试环境的样本数据。
适用场景
当你需要测试数据、模拟数据集或样本记录,用于开发、演示或填充测试环境时使用。适用于可以指定列结构、行数和输出格式的场景。
运行要求
该技能不附带脚本,仅为说明文档。若执行其生成的 Python 脚本,需要 Python 运行时及标准库(csv、json、datetime、random)。

Dummy Dataset Generation

Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script). Creates executable scripts or direct data files for immediate use.

Use when: Creating test data, generating sample datasets, building realistic mock data for development, or populating test environments.

Arguments:

  • $PRODUCT: The product or system name
  • $DATASET_TYPE: Type of data (e.g., customer feedback, transactions, user profiles)
  • $ROWS: Number of rows to generate (default: 100)
  • $COLUMNS: Specific columns or fields to include
  • $FORMAT: Output format (CSV, JSON, SQL, Python script)
  • $CONSTRAINTS: Additional constraints or business rules

Step-by-Step Process

  1. Identify dataset type - Understand the data domain
  2. Define column specifications - Names, data types, and value ranges
  3. Determine row count - How many sample records needed
  4. Select output format - CSV, JSON, SQL INSERT, or Python script
  5. Apply realistic patterns - Ensure data looks authentic and valid
  6. Add business constraints - Respect business logic and relationships
  7. Generate or script data - Create executable output
  8. Validate output - Ensure data quality and completeness

Template: Python Script Output

python
import csvimport jsonfrom datetime import datetime, timedeltaimport random
# ConfigurationROWS = $ROWSFILENAME = "$DATASET_TYPE.csv"
# Column definitions with realistic value generatorscolumns = {    "id": "auto-increment",    "name": "first_last_name",    "email": "email",    "created_at": "timestamp",    # Add more columns...}
def generate_dataset():    """Generate realistic dummy dataset"""    data = []    for i in range(1, ROWS + 1):        record = {            "id": f"U{i:06d}",            # Generate values based on column definitions        }        data.append(record)    return data
def save_as_csv(data, filename):    """Save dataset as CSV"""    with open(filename, 'w', newline='') as f:        writer = csv.DictWriter(f, fieldnames=data[0].keys())        writer.writeheader()        writer.writerows(data)
if __name__ == "__main__":    dataset = generate_dataset()    save_as_csv(dataset, FILENAME)    print(f"Generated {len(dataset)} records in {FILENAME}")

Example Dataset Specification

Dataset Type: Customer Feedback

Columns:

  • feedback_id (auto-increment, U001, U002...)
  • customer_name (realistic names)
  • email (valid email format)
  • feedback_date (dates last 90 days)
  • rating (1-5 stars)
  • category (Bug, Feature Request, Complaint, Praise)
  • text (realistic feedback)
  • product (electronics, clothing, home)

Constraints:

  • Ratings skewed: 40% 5-star, 30% 4-star, 20% 3-star, 10% 1-2 star
  • Bug category only with ratings 1-3
  • Feature requests only with ratings 3-5
  • Email domains realistic (gmail, yahoo, company.com)

Output Deliverables

  • Ready-to-execute Python script OR direct data file
  • CSV file with proper headers and formatting
  • JSON file with valid structure and types
  • SQL INSERT statements for database population
  • Data validation and constraint compliance
  • Realistic, business-appropriate values
  • Documentation of data generation logic
  • Quick-start instructions for using the dataset

Output Formats

CSV: Flat tabular format, easy to import into spreadsheets and databases

JSON: Nested structure, ideal for APIs and NoSQL databases

SQL: INSERT statements, directly executable on relational databases

Python Script: Executable generator for custom or large datasets

来源与署名

来源:phuryn/pm-skills位于pm-execution/skills/dummy-dataset提交8607e3b

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架