F1 Test Drive
Applicability (required before setup)
This is the canonical F1 applicability policy. Inspect the actual diff against the PR base and identify the behavior that needs validation before starting F1, selecting a fixture, or creating a report. Task size, issue labels, filenames, and a generic request to verify/ship are not sufficient reasons to run F1.
- Required: product/runtime changes exercised by F1 (issue tracking, authentication, routing, runner/session lifecycle, activity rendering), and functional changes to the F1 harness itself. Choose scenarios with assertions that exercise the changed behavior; a generic health check or unrelated fixture does not establish correctness.
- Not applicable: documentation, agent instructions, CI/release tooling, installer/build/packaging metadata, or other changes with no relevant product workflow behavior. Run appropriate documentation, script, unit, integration, build, package inspection, or isolated install smoke checks instead. Do not run F1 or create, attach, or commit an F1 test-drive report. A brief PR note stating why F1 is not applicable and what checks passed is enough.
- Mixed changes: exercise only the relevant functional portion with F1; validate other portions with their targeted checks. A docs, infrastructure, or metadata path is never an exemption for a real runtime behavior change (including dependencies, generated prompts, or installed defaults).
- Actual releases: assess the entire released payload against the previous
release, not merely the version-bump PR. Runtime-bearing payloads retain
relevant F1 validation. Payloads with no F1-covered behavior change use the
non-F1 release verification described in
apps/cli/RELEASING.md. Editing release tooling is not itself preparing or publishing a functional release.
If F1 is not applicable, stop this protocol before setup and reporting. If it is applicable but blocked, report the missing coverage and blocker honestly; do not relabel it as not applicable or substitute an unrelated passing drive. Preserve historical reports.
Mission (when applicable)
Validate the changed workflow through the relevant issue-tracker, EdgeWorker, and activity output paths. Name the scenario, expected behavior, and assertions before setup. Use the phases below as building blocks; include only paths needed for that scenario, plus setup and cleanup.
Test Drive Protocol
Phase 1: Setup
-
Create a fresh test repository (if needed):
-
Start F1 server:
-
Verify server health:
Phase 2: Issue-Tracker Verification
-
Create test issue:
-
Verify issue ID and issue creation response.
Phase 3: EdgeWorker Verification
-
Start agent session:
-
Monitor activities:
-
Verify:
- session started
- activities appear
- agent is processing issue
Phase 3.5: Slack Chat Session Verification (optional)
Use when validating the Slack → ChatSessionHandler → ClaudeRunner path. F1 exposes a test-only endpoint /cli/dispatch-chat that injects a synthetic app_mention event without going through Slack signature verification (SlackChatAdapter no-ops Slack API calls when slackBotToken is undefined).
-
Dispatch a synthetic chat event:
The response contains a
threadKeyof the form<channel>:<ts>. Reuse the same--thread-tsto address the same chat thread on subsequent dispatches. -
Verify shared auto-memory wiring:
- The chat workspace exists at
<cyrusHome>/slack-workspaces/<sanitized-threadKey>/. - The shared auto-memory directory exists (or is lazily creatable) at
<cyrusHome>/slack-memory/. - The
claude_query_optionsevent emitted byClaudeRunnercarriescqo.settingsAutoMemoryDirectory=<cyrusHome>/slack-memory.
- The chat workspace exists at
-
Verify per-thread workspace isolation alongside shared memory:
- Dispatch a second event in a different channel/thread.
- Confirm a separate
slack-workspaces/<other-thread-key>/directory exists (workspaces remain isolated). - Confirm both dispatches' telemetry resolve to the same
slack-memorypath (memory is shared).
Phase 4: Renderer Verification
-
Validate activity payload quality:
- expected types (for example
thought,action,response) - timestamps present
- content well-formed and readable
- expected types (for example
-
Validate pagination behavior:
Phase 5: Cleanup
-
Stop active session:
-
Stop background server process.
Reporting Format
Only after an applicable F1 drive, write a report under apps/f1/test-drives/.
Record the tested commit, changed behavior, commands, assertions, results, and
limitations. Adapt this template to the relevant scenario; omit unrelated
checklists. Never create a placeholder or “not applicable” F1 report:
Pass/Fail Criteria
Pass only when the changed-behavior assertions pass, along with the applicable workflow checks below:
- Server starts
- Issue created successfully
- Session starts and activities appear
- Activity payloads are coherent
- Session stops cleanly
- No unhandled errors
Fail when:
- server startup fails
- issue creation fails
- session does not start
- no activities after reasonable wait
- malformed activity data
- unhandled exceptions
Important Notes
- Prefer fixed port
3600unless already in use. - Use fresh test repos per drive.
- Preserve failed state when debugging.
- For functional runner/harness changes, validate the affected path end-to-end before merge, as required by the applicability policy above.
Multi-Harness Note
This skill is intentionally harness-agnostic:
- Claude subagents can call this skill.
- Codex/OpenCode workflows can reference the same skill content.
- Harness-specific adapters should be thin wrappers around this canonical skill.


