docs/clients/client.mdx
import { VersionBadge } from '/snippets/version-badge.mdx'
<VersionBadge version="2.0.0" />The fastmcp.Client class provides a programmatic interface for interacting with any MCP server. It handles protocol details and connection management automatically, letting you focus on the operations you want to perform.
The FastMCP Client is designed for deterministic, controlled interactions rather than autonomous behavior, making it ideal for testing MCP servers during development, building deterministic applications that need reliable MCP interactions, and creating the foundation for agentic or LLM-based clients with structured, type-safe operations.
<Note> This is a programmatic client that requires explicit function calls and provides direct control over all MCP operations. Use it as a building block for higher-level systems. </Note>You provide a server source and the client automatically infers the appropriate transport mechanism.
import asyncio
from fastmcp import Client, FastMCP
# In-memory server (ideal for testing)
server = FastMCP("TestServer")
client = Client(server)
# HTTP server
client = Client("https://example.com/mcp")
# Local Python script
client = Client("my_mcp_server.py")
async def main():
async with client:
# List available operations
tools = await client.list_tools()
resources = await client.list_resources()
prompts = await client.list_prompts()
# Execute operations
result = await client.call_tool("example_tool", {"param": "value"})
print(result)
asyncio.run(main())
All client operations require using the async with context manager for proper connection lifecycle management.
The client automatically selects a transport based on what you pass to it, but different transports have different characteristics that matter for your use case.
In-memory transport connects directly to a FastMCP server instance within the same Python process. Use this for testing and development where you want to eliminate subprocess and network complexity. The server shares your process's environment and memory space.
from fastmcp import Client, FastMCP
server = FastMCP("TestServer")
client = Client(server) # In-memory, no network or subprocess
STDIO transport launches a server as a subprocess and communicates through stdin/stdout pipes. This is the standard mechanism used by desktop clients like Claude Desktop. By default, the subprocess receives the MCP SDK's default environment; pass an explicit transport when you need to add environment variables, set a working directory, or control process reuse.
from fastmcp import Client
from fastmcp.client.transports import PythonStdioTransport
# Simple inference from file path
client = Client("my_server.py")
# With explicit environment configuration
transport = PythonStdioTransport(
"my_server.py",
env={"API_KEY": "secret"},
)
client = Client(transport)
HTTP transport connects to servers running as web services. Use this for production deployments where the server runs independently and manages its own lifecycle.
from fastmcp import Client
client = Client("https://api.example.com/mcp")
See Transports for detailed configuration options including authentication headers, session persistence, and multi-server configurations.
Create clients from MCP configuration dictionaries, which can include multiple servers. While there is no official standard for MCP configuration format, FastMCP follows established conventions used by tools like Claude Desktop.
config = {
"mcpServers": {
"weather": {
"url": "https://weather-api.example.com/mcp"
},
"assistant": {
"command": "python",
"args": ["./assistant_server.py"]
}
}
}
client = Client(config)
async with client:
# Tools are prefixed with server names
weather_data = await client.call_tool("weather_get_forecast", {"city": "London"})
response = await client.call_tool("assistant_answer_question", {"question": "What's the capital of France?"})
# Resources use prefixed URIs
icons = await client.read_resource("weather://weather/icons/sunny")
The client uses context managers for connection management. When you enter the context, the client establishes a connection and performs an MCP initialization handshake with the server. This handshake exchanges capabilities, server metadata, and instructions.
from fastmcp import Client, FastMCP
mcp = FastMCP(name="MyServer", instructions="Use the greet tool to say hello!")
@mcp.tool
def greet(name: str) -> str:
"""Greet a user by name."""
return f"Hello, {name}!"
async with Client(mcp) as client:
# Initialization already happened automatically
print(f"Server: {client.initialize_result.server_info.name}")
print(f"Instructions: {client.initialize_result.instructions}")
print(f"Capabilities: {client.initialize_result.capabilities.tools}")
For advanced scenarios where you need precise control over when initialization happens, disable automatic initialization and call initialize() manually:
from fastmcp import Client
client = Client("my_mcp_server.py", auto_initialize=False)
async with client:
# Connection established, but not initialized yet
print(f"Connected: {client.is_connected()}")
print(f"Initialized: {client.initialize_result is not None}") # False
# Initialize manually with custom timeout
result = await client.initialize(timeout=10.0)
print(f"Server: {result.server_info.name}")
# Now ready for operations
tools = await client.list_tools()
MCP has two protocol eras: the original legacy era, which begins every connection with an initialize handshake, and the modern era (protocol version 2026-07-28 and later), which a client discovers by probing the server's server/discover endpoint. The mode parameter controls which era the client negotiates when it connects.
By default, mode="auto". The client probes server/discover and adopts the modern protocol when the server responds; for any server that is not positive evidence of modern support, it falls back to the legacy handshake. This makes the default safe against a mixed fleet of legacy and modern servers.
from fastmcp import Client
# Negotiate the newest era the server supports (the default)
client = Client("https://example.com/mcp", mode="auto")
Set mode="legacy" to force the initialize handshake. This behaves identically to earlier FastMCP versions and is the opt-out if a server misbehaves under discovery or you need the legacy initialize result object.
client = Client("https://example.com/mcp", mode="legacy")
Legacy mode is also what you need for the capabilities that depend on a live session between client and server. The handshake opens a persistent back-channel the server can push requests down, and the modern era removed it. Pin mode="legacy" when your code relies on any of these:
task=Trueclient.ping() and transport.get_session_id()A FastMCP server serves both eras, so a default client negotiates the modern one and these raise an era-specific error. Pinning the handshake restores them.
You can also pin a specific modern protocol version to adopt it directly, without a discovery probe:
client = Client("https://example.com/mcp", mode="2026-07-28")
Once connected, the negotiated version and the server's advertised capabilities are available as properties. Both are populated regardless of which era was negotiated, and both are None while the client is disconnected.
async with Client("https://example.com/mcp", mode="auto") as client:
print(client.protocol_version) # e.g. "2026-07-28"
print(client.server_capabilities) # ServerCapabilities | None
The SSE transport is legacy-only — it cannot carry the sessionless modern era — so a client connecting over SSE always negotiates the legacy handshake, even under mode="auto". A multi-server config (MCPConfigTransport with more than one server) is likewise legacy-only, because it mounts each backend behind a legacy-era proxy; a single-server config mirrors its one backend transport's era.
</Note>
The client can cache the results of list_tools, list_resources, list_prompts, and read_resource so that repeated calls avoid a network round-trip. Caching is opt-in and honors the server's own cache hints, so it only takes effect against modern-era servers that advertise them — a cache is inert on a legacy connection.
Enable the default in-memory cache by passing cache=True. It respects the ttlMs and cacheScope hints the server attaches to each response.
from fastmcp import Client
client = Client("https://example.com/mcp", mode="auto", cache=True)
async with client:
tools = await client.list_tools() # fetched from the server
tools = await client.list_tools() # served from the cache
The default (cache=None) and cache=False both disable caching. For control over the store, TTL, or partitioning, pass a CacheConfig. A custom config requires a target_id, since in-memory FastMCP transports expose no server URL to derive a shared-store identity from.
from fastmcp import Client
from mcp.client.caching import CacheConfig
config = CacheConfig(target_id="weather-api", default_ttl_ms=60_000)
client = Client("https://example.com/mcp", mode="auto", cache=config)
The high-level list_tools, list_resources, list_prompts, and read_resource methods always use the cache when one is configured. To override the behavior for a single call, use the lower-level *_mcp variants, which accept a cache_mode argument: "use" (the default) serves and stores, "refresh" stores a fresh result without serving a cached one, and "bypass" skips the cache entirely.
async with client:
fresh = await client.list_tools_mcp(cache_mode="refresh")
The default cache lives in each client's process. To share cached responses across a fleet — a set of proxy replicas backed by one Redis, for example — pass a KeyValueResponseCacheStore, FastMCP's adapter over the same AsyncKeyValue key-value abstraction the event store and OAuth proxy use. It accepts any compatible backend (memory, Redis, and more).
A shared store mingles responses from different principals, so it requires an explicit partition that isolates them. Derive the partition from a verified credential — never from request data or the server URL — and construct a new client when the principal changes. Only responses the server marks "public" are ever served across partitions.
from fastmcp import Client
from fastmcp.client.caching import KeyValueResponseCacheStore
from mcp.client.caching import CacheConfig
from key_value.aio.stores.redis import RedisStore
backend = RedisStore(url="redis://localhost")
store = KeyValueResponseCacheStore(storage=backend)
config = CacheConfig(store=store, partition="tenant-a", target_id="weather-api")
client = Client("https://example.com/mcp", mode="auto", cache=config)
The adapter serializes each result through a type-tagged envelope validated against an allowlist of cacheable result models, so a value naming an unknown type is treated as a cache miss rather than deserialized blindly. Each store instance owns its own collection namespace; clear() affects only that namespace, never another tenant's entries.
Client extensions (SEP-2133) are the advanced mechanism a client uses to opt into vendor capabilities that live outside the core protocol. An extension is a ClientExtension instance that bundles three things: a capability advertisement the server can read, one or more result claims that let the client parse extra tools/call result shapes, and notification bindings that observe server notifications the core protocol doesn't define. Pass a sequence of them to extensions=.
from fastmcp import Client
from myproject.extensions import AppsExtension
client = Client("https://example.com/mcp", extensions=[AppsExtension()])
Each extension's contributions are threaded into the underlying session. Notification bindings compose with FastMCP's own internal task-status binding rather than replacing it, so an extension that observes a custom notification and FastMCP's task tracking both work on the same connection. When a tool returns a shape an extension claims, client.call_tool() resolves it transparently through the owning claim's resolver and hands you back an ordinary result. Result claims and their advertisements are honored only on modern-era connections, so they are inert on a legacy handshake.
For the rare case where you need to register additional result claims against an extension that is already advertised, pass them through result_claims=, keyed by the extension's identifier. Prefer declaring claims on the extension itself; this parameter merges extra claims with an extension's own.
client = Client(
"https://example.com/mcp",
extensions=[AppsExtension()],
result_claims={"example.com/apps": [extra_claim]},
)
FastMCP clients interact with three types of server components.
Tools are server-side functions that the client can execute with arguments. Call them with call_tool() and receive structured results.
async with client:
tools = await client.list_tools()
result = await client.call_tool("multiply", {"a": 5, "b": 3})
print(result.data) # 15
See Tools for detailed documentation including version selection, error handling, and structured output.
Resources are data sources that the client can read, either static or templated. Access them with read_resource() using URIs.
async with client:
resources = await client.list_resources()
content = await client.read_resource("file:///config/settings.json")
print(content[0].text)
See Resources for detailed documentation including templates and binary content.
Prompts are reusable message templates that can accept arguments. Retrieve rendered prompts with get_prompt().
async with client:
prompts = await client.list_prompts()
messages = await client.get_prompt("analyze_data", {"data": [1, 2, 3]})
print(messages.messages)
See Prompts for detailed documentation including argument serialization.
The client supports callback handlers for advanced server interactions. These let you respond to server-initiated requests and receive notifications.
Sampling, elicitation, and roots are all server-initiated, so they belong to the handshake era described under protocol negotiation. A default client negotiates the newest era both peers share, where the server has no back-channel to push those requests down, so an example that exercises them pins mode="legacy". Logging and progress arrive as notifications on the response stream and work in either era.
from fastmcp import Client
from fastmcp.client.logging import LogMessage
async def log_handler(message: LogMessage):
print(f"Server log: {message.data}")
async def progress_handler(progress: float, total: float | None, message: str | None):
print(f"Progress: {progress}/{total} - {message}")
async def sampling_handler(messages, params, context):
# Integrate with your LLM service here
return "Generated response"
client = Client(
"my_mcp_server.py",
mode="legacy",
log_handler=log_handler,
progress_handler=progress_handler,
sampling_handler=sampling_handler,
timeout=30.0
)
Each handler type has its own documentation: