Data Throughput Accelerator

作者 affaan-mef648e01899bMIT275K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库3天前更新

Diagnose and accelerate large data movement — ingestion, backfill, export, ETL, warehouse loading, manifest catch-up, and table synchronization — by isolating the true bottleneck, benchmarking variants, and codifying the fastest path with a hard accounting block proving rows and timestamps cohere. Use when a pipeline or backfill is too slow and must get faster without losing data correctness.

AI 生成的概览

诊断大规模数据移动缓慢问题,并固化一条更快且经正确性验证的管道路径。

功能
引导智能体诊断缓慢的数据移动,例如数据摄取、回填、导出、ETL、数仓加载、清单追赶和表同步。它会区分抽取、传输、加载、转换和新鲜度等瓶颈,对批次大小、工作进程数、暂存形态等变体进行基准测试,并推广在计数与时间戳保持一致的前提下最快的那条路径。它会产出一条固化路径(CLI、定时任务、工作流或运行手册),以及一个硬性核算块,报告文件数、行数、剩余尾部、运行时长和正确性门禁。
适用场景
适用于管道、回填或同步任务过慢,需要在不损失数据正确性的前提下提速的场景。适合瓶颈不明确、需要先测量再优化的大规模数据移动工作。
运行要求
仅为说明性指令,不附带脚本。智能体需要 Read、Write、Edit、Bash、Grep 和 Glob 工具,并能访问被测量的源、目标和清单系统。

Data Throughput Accelerator

Use this skill when the bottleneck is moving, transforming, or saving lots of data. The goal is not just speed. The goal is faster correct data landing in the right place with proof.

First Distinction

Separate these before optimizing:

  • source extraction speed;
  • network transfer speed;
  • warehouse/load speed;
  • transform speed;
  • serving-table freshness;
  • live tail growth while the job runs.

A pipeline can be "fast" and still appear behind if new data arrives faster than the final catch-up window.

Fast Path Heuristics

  • Move compute to where the data already is.
  • Prefer warehouse-native scans, joins, and appends for large landed files.
  • Use manifests or checkpoints so completed files/partitions are skipped.
  • Use partitioning and clustering that match the read and append pattern.
  • Batch small files, requests, and writes.
  • Make writes idempotent through unique keys, manifests, or replaceable staging.
  • Keep raw, derived, and serving tables separately accountable.

Workflow

  1. Read the current source, target, and manifest contracts.
  2. Measure backlog: external files, manifest rows, raw rows, derived rows, min/max timestamps, and unprocessed counts.
  3. Run a safe catch-up or sample benchmark.
  4. Compare variants: batch size, worker count, warehouse SQL, file grouping, staging shape, and manifest update method.
  5. Promote only the fastest path that keeps counts and timestamps coherent.
  6. Codify the path as a CLI, scheduled job, workflow, or runbook.
  7. Rerun final accounting after the codified path executes.

Accounting Output

Use a hard accounting block:

text
Data throughput result:- Source files discovered: 294- Files processed this run: 294- Raw rows added: 9,683,598- Derived rows added: 8,917,585- Remaining tail: 24 files at readback time- Runtime: 38.7s- Correctness gate: manifest counts and table max timestamps match

Guardrails

  • Do not delete raw data to make a metric look better.
  • Do not skip failed files silently.
  • Do not mix historical backfill status with live-tail freshness.
  • Do not call a pipeline complete until the target tables and manifest agree.
  • For finance, healthcare, regulated, or customer-impacting data, preserve replay evidence and approval gates.

来源与署名

来源:affaan-m/ecc位于skills/data-throughput-accelerator提交ef648e0

许可证: MIT

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架