CrewAI Agent Design Guide
How to design effective agents with the right role, goal, backstory, tools, and configuration.
Verified against crewai 1.15.23 on 2026-10-01. Live-tested with real LLMs on 2026-10-01.
Exact API forms (imports, defaults, provider extras, structured output) are in the check-crewai-api skill, and tools and MCP are in connect-tools-and-mcp. Where this skill and those disagree, follow them.
The 80/20 Rule
Spend 80% of your effort on task design, 20% on agent design. A well-designed task elevates even a simple agent. But even the best agent cannot rescue a vague, poorly scoped task. Get the task right first (see the design-task skill), then refine the agent.
0. How Many Agents Do You Actually Need?
Default to ONE agent. Add more only when the task genuinely splits into work that requires:
- Different tools or permissions — e.g. one agent has Slack write access, another reads docs only.
- Different personas the LLM must clearly switch between — a writer's voice is not a researcher's voice.
- Different LLMs — a cheap model for mechanical steps, a stronger one for synthesis.
- Different guardrails or output schemas — separate agents make the contract per stage explicit.
DO NOT add an agent just because the workflow has multiple steps. A single agent can call multiple tools in sequence within one kickoff (search → scrape → summarize is one agent's loop), produce structured multi-section output in one response, and iterate via its own tool-use loop.
Cost calculus: every extra agent = at least one more LLM kickoff plus a context handoff. Splitting linear, single-persona work into multiple agents multiplies token cost and adds fragility for marginal quality wins.
Anti-patterns
❌ Three agents for what is one researcher's job:
✅ One researcher does the gathering loop; one writer synthesizes — two agents because the personas and LLMs genuinely differ:
❌ A Summarizer agent plus a Slack Sender agent (apps=["slack"]) to summarize a string and DM it.
✅ One Slack Reporter agent with apps=["slack"] and a task: "Write a 2-3 sentence executive summary at the top, then DM {recipient_email} the summary followed by the full body."
Heuristic: if two "agents" share the same persona, the same tool surface, and the same LLM, they are one agent with a longer task description.
Once you've decided "one agent is enough"
Use Agent.kickoff() directly inside a Flow method - no Crew, no Task ceremony. The Flow owns sequencing and state; each step is a single agent kickoff (Flow mechanics: build-flow skill; docs: https://docs.crewai.com/en/concepts/agents#direct-agent-interaction-with-kickoff).
Reach for Crew.kickoff() only when a step genuinely benefits from multi-agent collaboration (delegation, hierarchical management, parallel specialists feeding one synthesis), or when the agent needs something that only works inside a crew task: knowledge sources and max_execution_time (Section 2).
1. The Role-Goal-Backstory Framework
Every agent needs three things: who it is, what it wants, and why it's qualified. {placeholders} in all three are filled from crew.kickoff(inputs={...}).
Role — Who the Agent Is
The role defines the agent's area of expertise. Be specific, not generic.
The role directly shapes how the LLM reasons. A "Senior Data Researcher" will produce different output than a "Research Assistant" even with the same task.
Goal — What the Agent Wants
The goal is the agent's individual objective. It should be outcome-focused with quality standards.
Backstory — Why the Agent Is Qualified
The backstory establishes expertise, experience, values, and working style. It's the agent's "personality prompt."
Include: years/depth of experience, specific domain knowledge, working style and values ("always cites sources", "prefers concise output"), and the quality standards the agent holds itself to.
Leave out: implementation details (tools, models, config), task-specific instructions (those go in the task description), and personality traits that don't affect output quality.
2. Agent Configuration Reference
Agent silently ignores keyword arguments it does not know, so a misspelled or removed parameter does nothing and raises nothing. Check "name" in Agent.model_fields when in doubt.
Execution limits (measured on 1.15.23)
So max_execution_time turns a slow run into an error but does not stop a hang. To bound wall-clock time, set timeouts where the time is spent: LLM(model=..., timeout=60) and timeouts inside your tools, or run the kickoff in a subprocess you can kill.
Tuning max_iter: the default 25 is generous - most tasks finish in 3-8 iterations. Lower it to 10-15 to fail faster when tasks are well-defined. If an agent keeps hitting it, the task is too vague (fix the task, not the limit). Hitting it does not raise, so check outputs, not exit codes.
Other switches: respect_context_window=True summarizes the conversation and continues when the provider reports that the context length was exceeded; with False the run stops with SystemExit: Context length exceeded .... inject_date=True is worth it for time-sensitive tasks (research, news, scheduling): asked for today's date, an agent with it answered correctly, and one without it answered UNKNOWN.
Tools
- An agent with no tools will hallucinate data when asked to search, fetch, or read files - always provide tools for tasks that require external data.
- Prefer fewer, focused tools - too many tools confuses the agent.
- Agent-level tools are available for all tasks the agent performs.
Task(tools=[...])replaces them for that task. In a live test, an agent with a weather tool given a task with only a population tool could call only the population tool. - Prebuilt tools need their keys (
SerperDevTool()needsSERPER_API_KEY). Custom tools, thecrewai_toolsnames that really exist, caching and MCP servers: connect-tools-and-mcp.
LLM Selection
- With no
llm=, the agent uses envMODEL, thenMODEL_NAME, thenOPENAI_MODEL_NAME, elsegpt-4.1-mini. That is an OpenAI model, so kickoff fails withoutOPENAI_API_KEY. - Only OpenAI ships with core crewai.
anthropic/...needsuv add "crewai[anthropic]",gemini/...needscrewai[google-genai], and non-native providers needcrewai[litellm](full table in check-crewai-api). - Give mechanical agents a cheaper model.
function_calling_llm(a separate model for tool calls) is deprecated on bothAgentandCrew- set each agent'sllminstead.
Collaboration
Set allow_delegation=True only when the agent is part of a crew with other specialized agents, the task genuinely benefits from handing off subtasks, or you're using hierarchical process where the manager delegates. Warning: delegation without clear task boundaries leads to infinite loops or wasted iterations.
Planning (Plan-and-Execute Mode)
With a PlanningConfig, the agent first generates a plan (a list of steps), executes each step in its own multi-turn loop (capped by max_step_iterations), checks each result, then continues, replans, or finishes early.
To disable planning, omit planning_config. planning=True alone is shorthand for PlanningConfig(reasoning_effort="low", max_attempts=1). reasoning=True is deprecated (DeprecationWarning) - use planning_config.
When to enable: for autonomous loops where the agent picks its own steps and you want failure recovery (a coding agent that writes → runs → patches; a research agent that searches → scrapes → revises). Skip it for single-tool, single-purpose calls ("summarize this string", "post this Slack DM").
Cost: In one measured run, planning added 3-7 calls to a two-step task and wrapped the answer in prose. Measure on your own tasks before turning it on.
Code Execution
allow_code_execution and code_execution_mode are deprecated no-ops (allow_code_execution=True emits a DeprecationWarning), and CodeInterpreterTool no longer exists in crewai_tools. Give the agent a sandbox tool such as E2BPythonTool or DaytonaPythonTool instead (see connect-tools-and-mcp).
Agent Guardrails
A guardrail can also be a string, which is checked by an extra LLM call. Agent guardrails have two limits on 1.15.23, both verified with a real LLM:
- They run only on
Agent.kickoff(). When the same agent executes aTaskin aCrew, the agent guardrail is never called. - On crewai 1.15.23 the agent-level guardrail retry re-sent the original prompt without the feedback, and all three attempts failed the uppercase check above. Prefer a task-level guardrail, which passes the feedback. When retries run out,
kickoffraisesValueError: Agent's guardrail failed validation after N retries. Last error: ....
When output must be fixed and not just rejected, put the guardrail on the Task (Task(guardrail=..., guardrail_max_retries=...)). The error message is fed back there, and the same uppercase check passed on the second attempt. See design-task and check-crewai-api.
Knowledge Sources
Knowledge sources give agents domain-specific data via RAG. Use them when agents need to reference large documents, policies, or datasets. Two things to know:
- Agent knowledge is only queried when the agent runs a task inside a
Crew.Agent.kickoff()silently ignoresknowledge_sources. In a live test, the same agent answered "UNKNOWN" fromkickoff()and correctly from a one-task crew. - Without
embedder=, knowledge uses OpenAI embeddings. With noOPENAI_API_KEY, agent knowledge raisesValueError: Invalid Knowledge Configurationat crew kickoff. Crew-level knowledge only logsFailed to upsert documentsand runs without it.
Embedders that work without OpenAI, KnowledgeConfig, memory and storage paths: references/memory-and-knowledge.md [blocked].
3. YAML Configuration (Recommended)
Define agents in agents.yaml for clean separation of config and code. Create a YAML project with crewai create crew <name> --classic. Without --classic, crewai create crew starts an interactive wizard that writes a JSON project instead.
Then wire in crew.py:
Critical: The method name (def researcher) must match the YAML key (researcher:). Mismatch causes KeyError (verified: KeyError: 'researcher').
Verified live: llm and max_iter set in YAML are applied, {topic} in the role is filled at kickoff, and tools attach in Python. Keep tools, guardrail functions and Pydantic models in Python; YAML holds the text and scalar settings.
4. Agent.kickoff() — Direct Agent Execution
Use Agent.kickoff() when you need one agent with tools and reasoning, without crew overhead. This is the most common pattern in Flows. It does not use the agent's knowledge sources or max_execution_time (Section 2).
Note:
Agent.kickoff()returnsLiteAgentOutput- access structured output viaresult.pydantic. This differs fromllm.call(messages, response_model=Model), which returns the Pydantic object directly.Agent.kickoff(..., response_model=...)is aTypeError.
File inputs need an extra: uv add "crewai[file-processing]", then from crewai_files import FileInput and researcher.kickoff("Summarize this document", input_files={"document": FileInput(path="report.pdf")}). Without the extra, the import fails with ModuleNotFoundError: No module named 'crewai_files'.
Agent.kickoff() vs Crew.kickoff(): use Agent.kickoff() when each step is a distinct agent and a Flow controls sequencing (the Section 0 shape - verified live as a two-step Flow). Use Crew.kickoff() when multiple agents collaborate on related tasks within a single step, or the agent needs knowledge sources or max_execution_time.
Agents in Conversational Flow Routes
In conversational Flows (from crewai.flow import ConversationState), the Flow owns the chat lifecycle and route selection. Call agents inside route handlers for bounded tool-backed work: research, docs lookup, account actions, triage, drafting, or escalation prep. The mechanics (handle_turn, ConversationConfig, a tested example) are in the build-flow skill's conversational-flows reference.
Design implications:
- Keep the conversational
Flowresponsible for session id, message history, routing, trace finalization, and approvals. - Keep each agent narrow: one route, one tool surface, one job.
- Use
self.append_agent_result(name, result, visibility="private")for scratch work that should not enter canonical chat history. - Return the user-visible reply from the handler (or call
self.append_assistant_message(reply)) so the next turn has the assistant context. - Do not make a "chat agent" with every tool. Route first, then invoke a focused agent for the selected route.
5. Specialist vs Generalist Agents
Apply this section after you've decided you genuinely need multiple agents (Section 0). With one agent, the only question is how to design that agent.
When you do need multiple agents, prefer specialists. An agent that does one thing well outperforms one that does many things acceptably. Use a specialist when the task needs deep domain knowledge, quality matters more than speed, or the task is complex enough to benefit from focused expertise. A generalist is acceptable for simple tasks with clear instructions, for prototyping you'll specialize later, and for tasks that truly span several domains equally.
Instead of one "Content Writer" agent, create technical_writer (technical accuracy, code examples), copywriter (persuasive, audience-focused copy) and editor (grammar, consistency, style guide). Each has a narrow role, specific goal, and a backstory that reinforces that expertise.
6. Agent Interaction Patterns
Sequential (default): Researcher → Writer → Editor. Agents work one after another, and each receives prior outputs as context. Best for linear pipelines where each step builds on the last.
Hierarchical: a manager agent delegates and validates; task assignment is dynamic. Best for complex workflows where assignment depends on intermediate results.
Without manager_llm or manager_agent, Crew(...) raises a ValidationError.
Agent-to-agent delegation: with allow_delegation=True, an agent gets two tools, Delegate work to coworker and Ask question to coworker, that name the other crew members. Verified live: asked to get a sentence from the Writer, a lead agent called ask_question_to_coworker and returned the Writer's answer.
7. Common Agent Design Mistakes
8. Agent Design Checklist
Before deploying an agent, verify:
- Role is specific and domain-focused (not "Assistant" or "Helper")
- Goal includes desired outcome AND quality standards
- Backstory establishes expertise and working style
- Tools are assigned for any task requiring external data
- No excess tools — 3-5 per agent maximum
- max_iter is tuned for expected task complexity (10-15 for simple, 20-25 for complex)
- Timeouts are set where time is spent (LLM
timeout, tool timeouts);max_execution_timeis only a backstop for crew tasks - Guardrails for critical outputs are on the Task
- LLM is set explicitly, fits the task's complexity, and has its provider extra installed
- Knowledge has an explicit
embedder, and the agent runs inside a crew - Delegation is disabled unless genuinely needed
References
For deeper dives into specific topics, see:
- Custom Tools [blocked] — building your own tools with
@tooldecorator andBaseToolsubclass - Memory & Knowledge [blocked] - memory, knowledge sources, embedders that work without OpenAI, storage, scoping
For related skills:
- check-crewai-api - current imports, parameters, defaults and provider extras
- connect-tools-and-mcp - custom tools, real
crewai_toolsnames, caching, MCP servers - build-flow - Flow state, routing, persistence, conversational flows
- getting-started - project scaffolding, choosing the right abstraction
- design-task — task description/expected_output best practices, guardrails, structured output, dependencies
- test-crewai-project - testing agents and crews offline with a stub LLM
- ask-docs - query the live CrewAI docs for questions not covered by these skills



