cookbook/performance/comparison/README.md
Compares Agno against LangGraph, PydanticAI and CrewAI on the costs a framework imposes before any model is called: cold import and agent construction (one OpenAI model reference plus one function tool, the same shape for every framework).
Construction and import never call a provider, so these benchmarks run with a placeholder API key and no network.
These benchmarks need the performance environment, which holds all four frameworks next to an editable install of this checkout's agno:
./scripts/perf_setup.sh
.venvs/perfenv/bin/python cookbook/performance/comparison/run_all.py
Results land in cookbook/performance/results/comparison/summary.json
(with framework versions recorded) and are picked up automatically by
report.py.
The mocked model requests one tool call; the framework dispatches and executes the real function; a second model turn answers. Every variant asserts the tool actually executed. This is where Agno pays its deferred tool-schema extraction (the flip side of its construction number). CrewAI is excluded: with a custom model its tool use goes through a text-based action protocol whose format is internal to the framework version, so a mock would be testing the mock rather than the framework.
The conversation benchmarks come in matched configurations in both directions, so neither side's persistence philosophy is silently advantaged:
cache_session=True
over an in-memory database — the closest analogue of LangGraph's
always-cached InMemorySaver; PydanticAI passes message_history;
CrewAI chains tasks through Task.context. Nothing is durably
persisted by anyone.SqliteDb, LangGraph with
SqliteSaver; both serialize and write to a SQLite file every turn,
with a fresh database file per conversation. LangGraph's figure
includes one graph compile (the checkpointer binds at compile).
PydanticAI ships no persistence layer and CrewAI has no conversation
primitive, so neither appears in this row.Agno does not currently win the 25-turn benchmark in either configuration: its per-turn write path re-serializes conversation state that grows with length. The results are published as measured; the growth term is a known optimization target. Every variant asserts after the final turn that history actually accumulated, so a silently stateless conversation fails instead of producing a flattering number.
All conversation variants raise Agno's default history cap
(num_history_runs=3) so the full conversation stays in context, matching
the other frameworks, which carry uncapped history. CrewAI's conversation
rows use task-context chaining because it has no lightweight conversation
primitive, and its memory feature requires an embedding provider, which
would violate the no-network constraint.
The single-turn run benchmark replaces the model at each framework's own
model boundary: Agno via a Model subclass, LangGraph via langchain's
GenericFakeChatModel, PydanticAI via its public TestModel, CrewAI via a
BaseLLM subclass. Each framework skips its own provider wire-format work,
so every number is that framework's floor. CrewAI builds a fresh Task and
Crew per run because a crew kickoff is its unit of request execution; its
Agent is reused like the other frameworks' agents.
langgraph.prebuilt.create_react_agent,
which compiles a state graph per call. LangGraph 1.x deprecates this
entrypoint in favor of the separate langchain package's create_agent;
it remains the canonical langgraph-only API.