Data Enrichment

hubspot/agent-cli-skills/data-enrichment

作者 hubspota8eea0880838无许可证27 个星标收录于 2026年10月8日更新于 2026年10月8日仓库7天前更新

Match external CSV/JSONL records to CRM contacts (by email) or companies (by domain) and write enriched data back in one pass using `hubspot objects upsert`.

仅含说明Data & Analytics
AI 生成的概览

按邮箱或域名将 CSV/JSONL 记录 upsert 到 HubSpot 联系人/公司,支持预演与错误处理。

功能
该技能描述了一套工作流:按邮箱将外部 CSV 或 JSONL 记录匹配到 HubSpot CRM 联系人,或按域名匹配到公司,然后通过一次 upsert 写入补充数据。内容涵盖用 jq 重塑输入、以 dry-run 预览、提取 digest 与 confirm 值、执行 upsert,以及拆分每条记录的成功与错误输出。它还说明了使用 OR 搜索过滤器的只读替代方案(含五个过滤组上限),并给出破坏性操作的安全实践。
适用场景
当你需要以邮箱或域名等自然标识为键,从外部表格或 JSONL 文件批量创建或更新 CRM 记录时使用。它也适用于只想读取匹配的 CRM 记录而不写回,或需要检查并重试失败 upsert 行的场景。
运行要求
需要 hubspot CLI 和 jq,处理 CSV 输入时还需要 csvkit 或其他 CSV 转 JSONL 工具。它依赖 bulk-operations 技能提供的 JSONL 管道、dry-run/digest、历史记录和速率限制指引,并需要 HubSpot 凭据和网络访问。该技能不附带脚本,仅为说明文档。

Prereq: read bulk-operations/SKILL.md first — JSONL piping, dry-run/digest, history, and rate-limit hygiene live there. This skill is the upsert-by-natural-key workflow on top.

The core move: upsert, not search-then-create

hubspot objects upsert --type X --id-property <natural-key> reads JSONL on stdin and creates-or-updates each row in one CLI call (the CLI batches 100 rows per API request), keyed by a property (email for contacts, domain for companies). No race window, no branching. Do not loop search → empty? → create.

Per line in: {"id":"[email protected]","properties":{"firstname":"Jane","jobtitle":"VP"}} Per line out: {"id":"123","ok":true,"data":{...}} or {"ok":false,"error":{...}}. Order matches input. The CLI adds no fields of its own — data is the raw batch-upsert API result row.

CSV/JSONL → upsert stream

Reshape with jq, preview with --dry-run, then execute. upsert is irreversible, so the execute step re-pipes the SAME inputs plus the --digest/--confirm lifted from the preview line (upsert confirm = the row count, always). Always lowercase the natural key — CRM match is exact. Confirm available property names with hubspot properties list --type contacts; never hard-code a list. See bulk-operations/resources/json-patterns.md for reshape idioms.

bash
# CSV → JSONL (any tool); example using csvkitcsvjson external.csv | jq -c '.[]' > external.jsonl
# Previewcat external.jsonl \| jq -c '{id:(.email|ascii_downcase), properties:{firstname:.first, lastname:.last, jobtitle:.title, company:.company}}' \| hubspot objects upsert --type contacts --id-property email --dry-run \| tee /tmp/upsert.preview.jsonl
# Lift the digest + confirm (present at every row count; upsert confirm = the row count)digest=$(jq -r 'select(.digest != null) | .digest' /tmp/upsert.preview.jsonl)confirm=$(jq -r 'select(.digest != null) | .target.id' /tmp/upsert.preview.jsonl)
# Execute — same pipeline, plus --digest/--confirm, capture resultscat external.jsonl \| jq -c '{id:(.email|ascii_downcase), properties:{firstname:.first, lastname:.last, jobtitle:.title, company:.company}}' \| hubspot objects upsert --type contacts --id-property email --digest "$digest" --confirm "$confirm" \| tee /tmp/upsert.results.jsonl

Companies: swap --type companies --id-property domain and reshape with .domain|ascii_downcase as id.

Handle per-record OK / error output

Split with jq, inspect failure modes, retry just the failures after fixing the inputs:

bash
jq -c 'select(.ok==true)'  /tmp/upsert.results.jsonl > /tmp/upsert.ok.jsonljq -c 'select(.ok==false)' /tmp/upsert.results.jsonl > /tmp/upsert.failed.jsonljq -r '.error.status' /tmp/upsert.failed.jsonl | sort | uniq -c   # status → count

The CLI does not tag rows as created-vs-updated. If the batch-upsert API result row carries a new boolean, jq -r '.data.new' /tmp/upsert.ok.jsonl | sort | uniq -c splits them; otherwise compare .data.createdAt against .data.updatedAt.

429s: split the input and rerun smaller chunks (see bulk-operations rate-limit notes). 400s usually mean a bad property name or invalid enum value — fix the reshape, rerun the failed inputs.

Destructive-op safety

upsert itself is non-destructive, but write-back can clobber populated fields. Always --dry-run first and spot-check. For bulk delete or overwrite of existing data, follow the dry-run → digest → confirm flow in bulk-operations/SKILL.md. Recovery: hubspot history --since 1h.

Match without upsert: OR-search → update

When you only want to read matches (no write-back), or the natural key isn't a CRM property, use repeated --filter flags — each flag is one OR group.

Verified cap: 5 OR groups per call. 6+ returns 400 too many filterGroups (count: N, max allowed: 5). Chunk 5 at a time:

bash
# emails.txt: one lowercased email per linexargs -n5 < emails.txt | while read -r e1 e2 e3 e4 e5; do  args=()  for e in "$e1" "$e2" "$e3" "$e4" "$e5"; do [ -n "$e" ] && args+=(--filter "email=$e"); done  hubspot objects search --type contacts "${args[@]}" --properties email,firstname,companydone > /tmp/matches.jsonl
jq -c '{id, properties:{lifecyclestage:"marketingqualifiedlead"}}' /tmp/matches.jsonl \| hubspot objects update --type contacts --dry-run

For larger keyed enrichments, prefer upsert — one pipeline, no chunking math.

来源与署名

来源:hubspot/agent-cli-skills位于data-enrichment提交a8eea08

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架