Braintrust Tracing for Claude Code
Comprehensive guide to tracing Claude Code sessions in Braintrust, including sub-agent correlation.
Architecture Overview
Hook Event Flow
Trace Hierarchy
Sub-Agent Tracing: What Works and What Doesn't
What Doesn't Work
SessionStart doesn't receive the Task prompt.
We tried injecting trace context into Task prompts via PreToolUse:
But SessionStart only receives session metadata, not the modified prompt. The injected context is lost.
What DOES Work
Task spans in parent session contain everything:
agentId- identifier for the sub-agent runtotalTokens,totalToolUseCount- metricscontent- full agent response/summarytool_input.prompt- original task prompttool_input.subagent_type- agent type (e.g., "oracle")
SubagentStop hook receives the sub-agent's session_id:
- This equals the sub-agent's orphaned trace
root_span_id - Allows correlation between parent Task span and child trace
The Correlation Pattern
Current state: Sub-agents create orphaned traces (new root_span_id).
Correlation method:
- Query parent session's Task spans for agent metadata
- Match
agentIdor timing with orphaned traces - Sub-agent's
session_id= its trace'sroot_span_id
Future solution (not yet implemented):
This would link: Task.agentId + Task.child_session_id -> orphaned trace root_span_id
State Management
Per-Session State Files
Each session file contains:
Global State
Debugging Commands
Check if Tracing is Active
Query Braintrust Directly
Debug Hook Execution
Troubleshooting Checklist
-
No traces appearing:
- Check
TRACE_TO_BRAINTRUST=truein.claude/settings.local.json - Verify API key:
echo $BRAINTRUST_API_KEY - Check logs:
tail -20 ~/.claude/state/braintrust_hook.log
- Check
-
Sub-agents not linking:
- This is expected - sub-agents create orphaned traces
- Use
--agent-statsto find agent activity - Correlate via timing or
agentIdin parent Task span
-
Missing spans:
- Check
current_turn_span_idin session state - Ensure Stop hook runs (turn finalization)
- Look for "Failed to create" errors in log
- Check
-
State corruption:
- Remove session state:
rm ~/.claude/state/braintrust_sessions/*.json - Clear global cache:
rm ~/.claude/state/braintrust_global.json
- Remove session state:
Key Files
Environment Variables
Session Learnings
What We Learned About Sub-Agent Tracing (Dec 2025)
Attempted: Inject trace context via PreToolUse into Task prompts.
Result: Failed - SessionStart only receives session metadata, not the prompt.
Discovery: Task spans already contain rich sub-agent data:
metadata.agent_type- agent type fromsubagent_typemetadata.skill_name- skill from Skill tooltool_input- full prompt sent to agenttool_output- agent response
Current correlation path:
- Parent session Task span has
agentIdand timing - Sub-agent creates orphaned trace with
root_span_id = session_id - SubagentStop provides the sub-agent's
session_id - Manual correlation: match timing or use
session_idlink
Future work: Write child_session_id to Task span metadata from PostToolUse after SubagentStop.
What We Learned About Sub-Agent Correlation
The Problem
- Sub-agents spawned via Task tool create orphaned Braintrust traces
- Parent session has Task spans with
agentId, sub-agent has separatesession_id - No built-in link between them
What DOESN'T Work
1. Prompt injection via PreToolUse
SessionStart hook only receives session metadata (session_id, type, cwd), NOT the prompt. Injected trace context is never seen.
The hook receives:
No prompt field exists - context injection is impossible at SessionStart.
2. SubagentStop → PostToolUse file handoff
Race condition. These are independent async hooks with no timing guarantees:
- SubagentStop fires when sub-agent session ends
- PostToolUse (Task) fires when Task tool completes
- No ordering guarantee between them
- Writing to a correlation file creates a race
3. PreToolUse correlation files
SessionStart can't access the task_span_id because it has no context about which Task spawned it. PreToolUse modifies prompts but doesn't create a reliably accessible state file that SessionStart can find.
What DOES Work
Post-hoc matching for dataset building:
Parent session Task spans contain:
agentId- identifier for the sub-agent runtotalTokens,totalToolUseCount- aggregated metricscontent- full agent response/summarytool_input.prompt- original task prompttool_input.subagent_type- agent type (e.g., "oracle")- Start/end timestamps
Sub-agent sessions contain:
session_id(equals orphaned traceroot_span_id)- Start/end timestamps
- All internal spans and tool calls
Correlation strategy:
- Export parent session traces (query parent
root_span_id) - Export sub-agent traces (query all sessions created within parent's time window)
- Match by:
- Timing: Task span end ≈ sub-agent session end
- Metadata:
subagent_typefrom Task prompt - IDs: SubagentStop hook provides
session_id(can be captured and logged)
Architecture Insight
SessionStart input is intentionally minimal - it contains no prompt or tool context:
This design boundary prevents real-time correlation at hook time.
Recommendation
For building agent run datasets with sub-agent correlation:
- In-session logging: Capture SubagentStop
session_idin logs or state - Post-session export: Query Braintrust API for parent and sub-agent traces
- Offline correlation: Match traces by timing and metadata in a script
- Don't try real-time linking: Hooks don't have necessary context
Example script pattern:
This approach is reliable, testable, and doesn't require hooks to maintain implicit state.


