# mcp client composition an mcp client represents one client-server relationship. a host that uses several servers normally owns several clients; "multi-server client" can describe several different architectures that should not be collapsed into one abstraction. ## endpoint composition endpoint composition presents several upstream servers as one synthetic mcp endpoint. fastmcp's `Client(config)` does this through an in-process router and proxy clients. this is useful when the consumer needs one client/session-shaped interface, but the proxy chain must use one compatible protocol era. a legacy connection can receive server-initiated sampling or elicitation, while a modern connection uses result-based input rounds and has no equivalent back-channel. forwarding halfway between those interaction models is not protocol translation. therefore an aggregate proxy can choose a shared era, but it cannot preserve independently negotiated eras. mixed legacy and modern servers require independent client-server connections. ## raw session grouping the official python sdk provides `ClientSessionGroup`. it manages raw `ClientSession` objects, aggregates tools, resources, and prompts, detects collisions, routes tool calls, and supports dynamic membership. in the sdk 2.1.1 implementation, `connect_to_server()` creates a session and calls `initialize()`. a real-world test against a stdio server and two remote http servers consequently negotiated the handshake era (`2025-11-25`) for all three. mixed eras still work with the upstream group when connection ownership is separated from aggregation: 1. connect one fastmcp `Client` per server, allowing each to negotiate independently; 2. register each active `client.session` with `ClientSessionGroup.connect_with_session()`; 3. use the sdk group for namespaced discovery and raw tool routing. this produced simultaneous `2025-11-25` and `2026-07-28` connections and routed real tool calls successfully. calls made through the group go directly through the registered raw sessions, not through the higher-level fastmcp `Client.call_tool()` behavior. ## application tool composition an agent framework usually needs native framework tools rather than one synthetic mcp endpoint. the natural unit is a one-server adapter: 1. a client discovers one server's mcp tool specifications; 2. the adapter converts each specification into a native tool; 3. each generated callable retains the originating client and upstream tool name; 4. a framework-level collection prefixes names and combines those native tools. because the callable already closes over its client, no additional mcp group is required for routing. the framework owns model-facing conversion, error normalization, metadata, middleware, and agent integration; the mcp client owns protocol negotiation and transport behavior. ## connection lifecycle transient and persistent usage can share the same generated callable when client contexts are reference counted: ```python async def call_tool(**arguments): async with client: return await client.call_tool(tool_name, arguments) ``` without an outer context, discovery or invocation opens and closes the selected connection. with an outer `async with client`, the inner context adds and releases a reference while leaving the warm connection alive. a higher-level adapter can therefore preserve existing stateless behavior and offer persistent connections without maintaining a second connection counter. caller-supplied clients remain important. handlers, authentication, caching, tracing, roots, and protocol mode belong to the client-server relationship and should not be flattened into global multi-server policy. ## choosing the boundary use the narrowest composition that matches the consumer: - use a proxy/router when several servers must appear as one mcp endpoint; - use `ClientSessionGroup` when raw sdk session aggregation is the desired interface; - use one client-backed adapter per server when producing framework-native tools; - add another shared group abstraction only after a concrete consumer demonstrates behavior that neither the sdk group nor downstream composition can express. retaining calls on a higher-level client surface can be a real distinction: it may preserve result parsing, multi-round tool handling, caching, tracing, or reference-counted lifecycle. it is not by itself evidence that another public group type is needed. write the downstream adapter first and measure how much non-policy code remains. ## design lessons - inspect upstream primitives before designing a parallel abstraction; - distinguish endpoint composition from application-side collection; - prototype the actual consumer, not only transport and routing mechanics; - test mixed protocol eras against real stdio and http servers; - treat namespacing, partial availability, refresh policy, and lifecycle ownership as product choices rather than incidental plumbing; - keep an experimental abstraction draft until a consumer proves that it removes meaningful work. ## sources - [MCP architecture, protocol revision 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28/architecture) - [MCP Python SDK session groups](https://py.sdk.modelcontextprotocol.io/client/session-groups/) - [MCP Python SDK `ClientSessionGroup`](https://github.com/modelcontextprotocol/python-sdk/blob/main/src/mcp/client/session_group.py), reviewed 2026-08-26 - [FastMCP PR #4893: modern protocol in proxy-backed multi-server clients](https://github.com/PrefectHQ/fastmcp/pull/4893) - [FastMCP draft PR #4904: independent client-group experiment](https://github.com/PrefectHQ/fastmcp/pull/4904) - [LangChain PR #39922: one-server FastMCP-backed adapter](https://github.com/langchain-ai/langchain/pull/39922)