E2E Behavior Validation for Frontend Modifications
Core Principle: Test Product Behavior, Not UI States
CRITICAL: Tests must verify that product features WORK correctly, not just that UI elements render.
What NOT to test (UI States):
- ❌ "Dropdown opens when clicked"
- ❌ "Modal appears after button click"
- ❌ "Loading spinner shows during request"
- ❌ "Form fields are visible"
- ❌ "Sidebar collapses"
What TO test (Product Behavior):
- ✅ "Selecting an LLM provider configures the agent to use that provider"
- ✅ "Creating a new agent persists it and shows in the agents list"
- ✅ "Running a tool with parameters returns the expected output"
- ✅ "Chat messages stream correctly and maintain conversation context"
- ✅ "Workflow execution triggers tools in the correct order"
BDD Structure (REQUIRED)
Every E2E spec MUST follow the same BDD shape as the MSW tests. In packages/playground, e2e-bdd/test-needs-when-describe enforces this shape.
The structure has exactly three levels:
- Outer
test.describe= the unit under test (one page or feature per file). - Inner
test.describe('when …')= exactly ONE precondition. The title MUST start withwhen. - Each
test= exactly ONE observable outcome.
Rules:
- One outer
test.describeper file naming the unit. - Every leaf
testlives inside atest.describe('when …')precondition group. No top-level flattest(). - Split a multi-assertion
test()only where assertions represent distinct outcomes; keep tightly-coupled assertions that prove a single outcome together. Never drop an assertion. - Place
beforeEach/afterEachin the narrowestdescribescope that needs them.
Prerequisites
Requires Playwright MCP server. If the browser_navigate tool is unavailable, instruct the user to add it:
Step 1: Understand the Feature Intent
Before writing ANY test, answer these questions:
- What user problem does this feature solve?
- What is the expected outcome when the feature works correctly?
- What data flows through the system? (user input → API → state → UI)
- What should persist after page reload?
- What downstream effects should this action have?
Document these answers as comments in your test file.
Step 2: Build and Start
Verify server at http://localhost:4111
Step 3: Map Feature to Behavior Tests
Feature-to-Test Mapping Guide
Step 4: Write Behavior-Focused Tests
Test Structure Template
Behavior Test Patterns
Pattern 1: Configuration Affects Behavior
Pattern 2: Data Persistence
Pattern 3: Tool Execution Produces Correct Output
Pattern 4: Workflow Step Chaining
Pattern 5: Streaming Chat with Context
Pattern 6: Error Recovery
Step 5: Update Existing Tests
When a test file already exists:
- Read the existing tests to understand current coverage
- Identify if tests are UI-focused or behavior-focused
- Refactor UI-focused tests to verify behavior instead:
Refactoring Example
BEFORE (UI-focused):
AFTER (Behavior-focused + BDD nesting):
Step 6: Kitchen-Sink Fixtures for Behavior Testing
Fixtures should represent realistic scenarios, not just mock data:
Fixture Naming Convention
Fixture Content Requirements
Each fixture must define:
- Scenario description (what behavior it enables testing)
- Expected outcomes (what assertions should pass)
- Edge cases covered (error states, empty states, etc.)
Step 7: Run and Validate
Test Quality Checklist
Before considering tests complete, verify:
- Each test has a clear user story comment
- One outer
test.describenames the unit under test - Every
testis nested in atest.describe('when …')precondition block (no flat top-leveltest()) - Each
testasserts exactly ONE observable outcome - Tests verify OUTCOMES, not intermediate UI states
- Tests would FAIL if the feature broke (not just if UI changed)
- Persistence is verified via
page.reload()where applicable - Error scenarios are covered
- Tests use appropriate timeouts for async operations
- Fixtures represent realistic usage scenarios


