Goldsky Dataset Reference
Reference tables for blockchain datasets available in Turbo pipelines.
For quick dataset questions (e.g., "what dataset for Solana transfers?"), answer directly: identify the chain prefix (see Popular Chain Prefixes below), identify the dataset type (see Common Datasets), and return a YAML snippet like:
Tip: Verify a dataset exists with
goldsky dataset get <chain>.<dataset> --outputFormat json— returns the exact slug, version, schema, andisTurboOnly(false= usable by Turbo and Mirror;true= Turbo-only). To browse or search:goldsky dataset list --output json(optionally--group <chain>; the list is a slim index and omitsisTurboOnly— usedataset getfor that).goldsky turbo validate file.yamlvalidates a whole Turbo pipeline config — it is not a dataset-existence check (Turbo-specific, needs theturbobinary).⚠️ Plain
goldsky dataset listwith no flags is an interactive picker that only works in a real terminal — in an agent/CI/piped context it produces no output at all. Always pass--output json(or--group). On an older CLI without those flags, fall back to the public read-only endpoint:curl -s https://api.goldsky.com/api/public/datasets/v1.
Dataset Reference Files
Chain prefix and chain ID information is in the
data/folder.
Data location: data/ (relative to this skill's directory)
Dataset versions are not pinned in this skill. Version numbers drift as datasets are revised, so confirm the exact version with
goldsky dataset get <name> --outputFormat jsonrather than trusting a static list. The dataset-type and schema tables below are stable references; treat any version number in them as a starting point to validate, not ground truth.
Quick Reference
Common Datasets
Important: Use
raw_transactions, NOTtransactions
Popular Chain Prefixes
See data/chain-prefixes.json for complete list with chain IDs.
Common Dataset Types
EVM Chains
Important: Use
raw_transactions, NOTtransactions. Useraw_logs, NOTlogs(thoughlogsworks as an alias on some chains).There is no consumable
decoded_logssource dataset. For decoded contract events, source from<chain>.raw_logsand decode in a SQL transform via_gs_log_decode(_gs_fetch_abi(<explorer-url>, <source>), topics, data). See/turbo-transformsfor the full pattern.
Solana
Bitcoin
Stellar
All datasets use version 1.1.0:
Sui
NEAR
Starknet
Fogo
Dataset Schemas
Source: docs.goldsky.com. Do not use field names not listed here — run
goldsky dataset get <chain>.<dataset> --outputFormat jsonto inspect unknown schemas.
Solana
solana.transactions
No
from_addressorto_addresson Solana transactions — useaccountsarray instead.
solana.transactions_with_instructions
All fields from solana.transactions plus:
Instruction object fields: id, index, parent_index, block_slot, block_timestamp, block_hash, tx_signature, tx_fee, tx_index, program_id, data (base58), accounts (string[]), instruction_type
solana.instructions
solana.token_transfers
solana.native_balances
solana.blocks
solana.rewards
solana.token_balances
Schema not fully documented — do not guess field names. Inspect with
goldsky dataset get solana.token_balances --outputFormat json.
EVM Chains
<chain>.raw_logs / <chain>.logs
topicsis a comma-separated string, not an array. Topic 0 is the event signature hash.
<chain>.raw_transactions
L2 chains also include:
receipt_l1_fee,receipt_l1_gas_used,receipt_l1_gas_price,receipt_l1_fee_scalar
<chain>.blocks
<chain>.erc20_transfers
<chain>.erc721_transfers
Dataset Name Format
All datasets follow the pattern: <chain_prefix>.<dataset_type>
Examples:
ethereum.erc20_transfers- ERC-20 transfers on Ethereum mainnetbase.logs- All event logs on Basematic.blocks- Block data on Polygonsolana.token_transfers- SPL token transfers on Solana
Finding Dataset Versions
Datasets are versioned. To find available versions:
Common versions:
1.0.0- Initial version1.2.0- Enhanced schema (common for ERC-20 transfers)
When in doubt, use the latest version shown by goldsky dataset get <name> --outputFormat json.
Common Discovery Patterns
"I want to track USDC transfers on Base"
- Dataset:
base.erc20_transfers - Filter by contract address in your pipeline transform:
"I want all NFT activity on Ethereum"
Dataset: ethereum.erc721_transfers
"I want to monitor a specific smart contract"
- Source:
<chain>.raw_logs, filtered to the contract address in the source filter - To decode events: add a SQL transform calling
_gs_log_decode(_gs_fetch_abi(<explorer-url>, <source>), topics, data) AS decoded, then filter downstream byWHERE decoded.event_signature = '<EventName>(<types>)'. Never put topic0 hashes in the source filter.
"I need multi-chain data"
Use multiple sources in your pipeline:
Troubleshooting
Dataset not found
Fix:
- Check the chain prefix is correct (e.g.,
maticnotpolygon) - Check the dataset type exists (e.g.,
erc20_transfersnoterc20) - Run
goldsky dataset list --output jsonto see all available options
Chain not listed
If you can't find a chain in the tables above:
Some chains use non-obvious prefixes (e.g., Polygon uses matic).
Version mismatch
Fix: Check available versions:
Use a version that exists in the output.
Related
/turbo-builder— Interactive wizard to build pipelines using these datasets

