Data Enrichment

hubspot/agent-cli-skills/data-enrichment

作者 hubspota8eea0880838無授權條款27 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫7 天前更新

Match external CSV/JSONL records to CRM contacts (by email) or companies (by domain) and write enriched data back in one pass using `hubspot objects upsert`.

僅含說明Data & Analytics
AI 產生的概覽

依電子郵件或網域將 CSV/JSONL 記錄 upsert 至 HubSpot 聯絡人/公司,支援預演與錯誤處理。

功能
此技能說明一套工作流程:依電子郵件將外部 CSV 或 JSONL 記錄比對到 HubSpot CRM 聯絡人,或依網域比對到公司,再透過一次 upsert 寫回補充資料。內容涵蓋以 jq 重塑輸入、用 dry-run 預覽、取出 digest 與 confirm 值、執行 upsert,以及拆分每筆記錄的成功與錯誤輸出。它也說明使用 OR 搜尋篩選器的唯讀替代做法(含五個篩選群組上限),並提供破壞性操作的安全實務。
適用情境
當你需要以電子郵件或網域等自然識別鍵,從外部試算表或 JSONL 檔案批次建立或更新 CRM 記錄時使用。它也適合只想讀取相符的 CRM 記錄而不寫回,或需要檢查並重試失敗 upsert 資料列的情況。
執行需求
需要 hubspot CLI 與 jq,處理 CSV 輸入時還需要 csvkit 或其他 CSV 轉 JSONL 工具。它依賴 bulk-operations 技能提供的 JSONL 管線、dry-run/digest、歷史記錄與速率限制指引,並需要 HubSpot 憑證與網路存取。此技能未附帶指令碼,僅為說明文件。

Prereq: read bulk-operations/SKILL.md first — JSONL piping, dry-run/digest, history, and rate-limit hygiene live there. This skill is the upsert-by-natural-key workflow on top.

The core move: upsert, not search-then-create

hubspot objects upsert --type X --id-property <natural-key> reads JSONL on stdin and creates-or-updates each row in one CLI call (the CLI batches 100 rows per API request), keyed by a property (email for contacts, domain for companies). No race window, no branching. Do not loop search → empty? → create.

Per line in: {"id":"[email protected]","properties":{"firstname":"Jane","jobtitle":"VP"}} Per line out: {"id":"123","ok":true,"data":{...}} or {"ok":false,"error":{...}}. Order matches input. The CLI adds no fields of its own — data is the raw batch-upsert API result row.

CSV/JSONL → upsert stream

Reshape with jq, preview with --dry-run, then execute. upsert is irreversible, so the execute step re-pipes the SAME inputs plus the --digest/--confirm lifted from the preview line (upsert confirm = the row count, always). Always lowercase the natural key — CRM match is exact. Confirm available property names with hubspot properties list --type contacts; never hard-code a list. See bulk-operations/resources/json-patterns.md for reshape idioms.

bash
# CSV → JSONL (any tool); example using csvkitcsvjson external.csv | jq -c '.[]' > external.jsonl
# Previewcat external.jsonl \| jq -c '{id:(.email|ascii_downcase), properties:{firstname:.first, lastname:.last, jobtitle:.title, company:.company}}' \| hubspot objects upsert --type contacts --id-property email --dry-run \| tee /tmp/upsert.preview.jsonl
# Lift the digest + confirm (present at every row count; upsert confirm = the row count)digest=$(jq -r 'select(.digest != null) | .digest' /tmp/upsert.preview.jsonl)confirm=$(jq -r 'select(.digest != null) | .target.id' /tmp/upsert.preview.jsonl)
# Execute — same pipeline, plus --digest/--confirm, capture resultscat external.jsonl \| jq -c '{id:(.email|ascii_downcase), properties:{firstname:.first, lastname:.last, jobtitle:.title, company:.company}}' \| hubspot objects upsert --type contacts --id-property email --digest "$digest" --confirm "$confirm" \| tee /tmp/upsert.results.jsonl

Companies: swap --type companies --id-property domain and reshape with .domain|ascii_downcase as id.

Handle per-record OK / error output

Split with jq, inspect failure modes, retry just the failures after fixing the inputs:

bash
jq -c 'select(.ok==true)'  /tmp/upsert.results.jsonl > /tmp/upsert.ok.jsonljq -c 'select(.ok==false)' /tmp/upsert.results.jsonl > /tmp/upsert.failed.jsonljq -r '.error.status' /tmp/upsert.failed.jsonl | sort | uniq -c   # status → count

The CLI does not tag rows as created-vs-updated. If the batch-upsert API result row carries a new boolean, jq -r '.data.new' /tmp/upsert.ok.jsonl | sort | uniq -c splits them; otherwise compare .data.createdAt against .data.updatedAt.

429s: split the input and rerun smaller chunks (see bulk-operations rate-limit notes). 400s usually mean a bad property name or invalid enum value — fix the reshape, rerun the failed inputs.

Destructive-op safety

upsert itself is non-destructive, but write-back can clobber populated fields. Always --dry-run first and spot-check. For bulk delete or overwrite of existing data, follow the dry-run → digest → confirm flow in bulk-operations/SKILL.md. Recovery: hubspot history --since 1h.

Match without upsert: OR-search → update

When you only want to read matches (no write-back), or the natural key isn't a CRM property, use repeated --filter flags — each flag is one OR group.

Verified cap: 5 OR groups per call. 6+ returns 400 too many filterGroups (count: N, max allowed: 5). Chunk 5 at a time:

bash
# emails.txt: one lowercased email per linexargs -n5 < emails.txt | while read -r e1 e2 e3 e4 e5; do  args=()  for e in "$e1" "$e2" "$e3" "$e4" "$e5"; do [ -n "$e" ] && args+=(--filter "email=$e"); done  hubspot objects search --type contacts "${args[@]}" --properties email,firstname,companydone > /tmp/matches.jsonl
jq -c '{id, properties:{lifecyclestage:"marketingqualifiedlead"}}' /tmp/matches.jsonl \| hubspot objects update --type contacts --dry-run

For larger keyed enrichments, prefer upsert — one pipeline, no chunking math.

來源與署名

來源:hubspot/agent-cli-skills位於data-enrichment提交a8eea08

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架