Creating Evaluator Functions
Overview
This skill documents how to create evaluator functions in evaluators.ts for Output SDK workflows. Evaluators are used to assess quality, validate outputs, and provide confidence-scored judgments about workflow results.
When to Use This Skill
- Implementing quality assessment for workflow outputs
- Adding validation logic with confidence scores
- Creating LLM-powered content evaluation
- Building reusable evaluation components
File Organization
Option 1: Flat File (Default)
For smaller workflows, use a single evaluators.ts file:
Option 2: Folder-Based (Large workflows)
For larger workflows with many evaluators, use an evaluators/ folder:
Component Location Rules
Important: evaluator() calls MUST be in files containing 'evaluators' in the path:
src/workflows/my_workflow/evaluators.ts✓src/workflows/my_workflow/evaluators/quality.ts✓src/shared/evaluators/common_evaluators.ts✓src/workflows/my_workflow/helpers.ts✗ (cannot contain evaluator() calls)
Activity Isolation Constraints
Evaluators are Temporal activities with strict import rules to ensure deterministic replay.
Evaluators CAN import from:
- Local workflow files:
./utils.js,./types.js,./helpers.js - Local subdirectories:
./lib/helpers.js - Shared utilities:
../../shared/utils/*.js - Shared clients:
../../shared/clients/*.js - Shared services:
../../shared/services/*.js
Evaluators CANNOT import:
- Other evaluator files (activity isolation)
- Step files
- Workflow files
Example of WRONG imports:
Critical Import Patterns
Core Imports
LLM Client Import (for LLM-powered evaluators)
ES Module Imports
All imports MUST use .js extension:
Basic Structure
Required Properties
name (string)
Unique identifier for the evaluator. Use snake_case.
description (string)
Human-readable description of what the evaluator assesses.
inputSchema (Zod schema)
Schema for validating evaluator input.
fn (async function)
The evaluator execution function. Returns an evaluation result with value and confidence.
Result Types
EvaluationBooleanResult
Use for pass/fail or true/false evaluations:
EvaluationNumberResult
Use for numeric scores or ratings:
EvaluationStringResult
Use for categorical or text-based evaluations:
Result Properties
Simple Evaluator Examples
Boolean Evaluator - Content Validation
Boolean Evaluator - Pattern Detection
Number Evaluator - Quality Score
String Evaluator - Sentiment Classification
LLM-Powered Evaluator Examples
Note: Evaluators are self-contained components that don't share schemas across steps, so defining aiSdk.Output.object() schemas inline is acceptable here. For workflow steps that share schemas, define them in types.ts instead.
generateText arguments: prompt, promptDir, variables, tools, output, toolChoice, stopWhen, abortSignal.
Using generateText with aiSdk.Output.object() for Evaluation
LLM Boolean Evaluation
LLM String Evaluation - Content Classification
EvaluationResult with Feedback
Use the feedback field to provide actionable improvement suggestions alongside your evaluation result. Import EvaluationFeedback from @outputai/core to create feedback objects.
EvaluationFeedback Properties
Multi-Dimensional Evaluation
Use the dimensions field to nest EvaluationResult instances for sub-scores. Each dimension should use the name field to identify it.
Complete Example
Based on a real workflow evaluator file:
Best Practices
1. Use Appropriate Result Types
2. Provide Meaningful Confidence Scores
3. Include Reasoning for Transparency
4. Keep Evaluators Focused
5. Use Descriptive Names
6. Use Feedback for Actionable Improvements
7. Use Dimensions for Multi-Criteria Evaluation
Verification Checklist
-
evaluator,z, result types imported from@outputai/core -
generateTextandaiSdkimported from@outputai/llmif using LLM (not direct provider) - LLM output schemas use
.describe()instead of.min()/.max()onz.number() - All imports use
.jsextension - Named exports used for each evaluator
- Each evaluator has
name,description,inputSchema,fn - Evaluator name uses
snake_case - Returns appropriate result type (
EvaluationBooleanResult,EvaluationNumberResult, orEvaluationStringResult) - Confidence score between 0.0 and 1.0
- Evaluators only import allowed dependencies (local files, shared code)
- No imports of other evaluators, steps, or workflows
-
EvaluationFeedbackimported from@outputai/corewhen using feedback - Feedback objects include
issue,suggestion, andpriority - Dimensions use the
namefield to identify sub-evaluations
Related Skills
output-dev-workflow-function- Orchestrating evaluators in workflow.tsoutput-dev-step-function- Creating step functionsoutput-dev-types-file- Defining evaluator input schemasoutput-dev-prompt-file- Creating prompt files for LLM-powered evaluatorsoutput-dev-folder-structure- Understanding project layoutoutput-eval-error-analysis— Identify what to evaluate before writing evaluators


