Data Throughput Accelerator

作者 affaan-mef648e01899bMIT275K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 天前更新

Diagnose and accelerate large data movement — ingestion, backfill, export, ETL, warehouse loading, manifest catch-up, and table synchronization — by isolating the true bottleneck, benchmarking variants, and codifying the fastest path with a hard accounting block proving rows and timestamps cohere. Use when a pipeline or backfill is too slow and must get faster without losing data correctness.

AI 產生的概覽

診斷大規模資料移動緩慢的問題,並固化出一條更快且經正確性驗證的管線路徑。

功能
引導代理診斷緩慢的資料移動,例如資料擷取、回填、匯出、ETL、資料倉儲載入、清單追趕和資料表同步。它會區分抽取、傳輸、載入、轉換和新鮮度等瓶頸,對批次大小、工作行程數、暫存形態等變體進行基準測試,並推廣在計數與時間戳保持一致的前提下最快的那條路徑。它會產出一條固化路徑(CLI、排程工作、工作流程或執行手冊),以及一個硬性核算區塊,報告檔案數、列數、剩餘尾端、執行時間和正確性閘門。
適用情境
適用於管線、回填或同步工作過慢,需要在不損失資料正確性的前提下加速的情況。適合瓶頸不明確、需要先量測再最佳化的大規模資料移動工作。
執行需求
僅為說明性指示,不附帶指令碼。代理需要 Read、Write、Edit、Bash、Grep 和 Glob 工具,並能存取被量測的來源、目標和清單系統。

Data Throughput Accelerator

Use this skill when the bottleneck is moving, transforming, or saving lots of data. The goal is not just speed. The goal is faster correct data landing in the right place with proof.

First Distinction

Separate these before optimizing:

  • source extraction speed;
  • network transfer speed;
  • warehouse/load speed;
  • transform speed;
  • serving-table freshness;
  • live tail growth while the job runs.

A pipeline can be "fast" and still appear behind if new data arrives faster than the final catch-up window.

Fast Path Heuristics

  • Move compute to where the data already is.
  • Prefer warehouse-native scans, joins, and appends for large landed files.
  • Use manifests or checkpoints so completed files/partitions are skipped.
  • Use partitioning and clustering that match the read and append pattern.
  • Batch small files, requests, and writes.
  • Make writes idempotent through unique keys, manifests, or replaceable staging.
  • Keep raw, derived, and serving tables separately accountable.

Workflow

  1. Read the current source, target, and manifest contracts.
  2. Measure backlog: external files, manifest rows, raw rows, derived rows, min/max timestamps, and unprocessed counts.
  3. Run a safe catch-up or sample benchmark.
  4. Compare variants: batch size, worker count, warehouse SQL, file grouping, staging shape, and manifest update method.
  5. Promote only the fastest path that keeps counts and timestamps coherent.
  6. Codify the path as a CLI, scheduled job, workflow, or runbook.
  7. Rerun final accounting after the codified path executes.

Accounting Output

Use a hard accounting block:

text
Data throughput result:- Source files discovered: 294- Files processed this run: 294- Raw rows added: 9,683,598- Derived rows added: 8,917,585- Remaining tail: 24 files at readback time- Runtime: 38.7s- Correctness gate: manifest counts and table max timestamps match

Guardrails

  • Do not delete raw data to make a metric look better.
  • Do not skip failed files silently.
  • Do not mix historical backfill status with live-tail freshness.
  • Do not call a pipeline complete until the target tables and manifest agree.
  • For finance, healthcare, regulated, or customer-impacting data, preserve replay evidence and approval gates.

來源與署名

來源:affaan-m/ecc位於skills/data-throughput-accelerator提交ef648e0

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架