Guide

Local vs remote MCP servers: which should you build

Most people ask this as a deployment question and answer it on performance grounds, which is the wrong axis. A local MCP server runs on your machine and inherits your credentials and your filesystem, which is exactly what you want for personal tooling and exactly what you cannot hand to a team. A remote one is a service you now own, with identity, authorisation, multi-tenancy and someone whose phone rings when it breaks.

What the spec actually says about transports right now

The Model Context Protocol specification defines two standard transport bindings: stdio and Streamable HTTP. The older HTTP+SSE transport from protocol version 2024-11-05 is formally deprecated, and the current spec is blunt about it: "New implementations SHOULD NOT adopt it; existing implementations SHOULD migrate to Streamable HTTP."

Neither transport is "the local one" or "the remote one" at the protocol level. The transports overview states that "Protocol semantics are identical on every transport", and the docs architecture page is equally clear that "MCP servers can execute locally or remotely."

What the docs give you instead is a statement about typical usage, and it is the sentence the whole decision hangs on: "Local MCP servers that use the STDIO transport typically serve a single MCP client, whereas remote MCP servers that use the Streamable HTTP transport will typically serve many MCP clients." The mechanics follow: on stdio "the client launches the MCP server as a subprocess" and speaks JSON-RPC over its standard streams; on Streamable HTTP it is an independent process behind a single HTTP endpoint that supports POST.

One warning, because several published guides are stale. Revision 2026-07-28 is current, and its Streamable HTTP page lists what changed:

  • "Removal of the GET stream endpoint. Removal of protocol-level sessions."
  • The same page adds that "Resumable SSE streams via Last-Event-ID are not supported".
  • The versioning page says flatly that there is no negotiation handshake.

A tutorial describing session IDs, an initialize negotiation or a long-lived GET stream is describing 2025-11-25 or earlier, which the spec classifies as Legacy. Even good material lands there: Microsoft's MCP for Beginners says in its README, "This curriculum is aligned with MCP Specification 2025-11-25 (the latest stable release)." The README flags the 2026-07-28 release candidate separately, but the transport material it teaches against is the legacy model.

Question one: who else needs this

If the answer is nobody, build local and stop reading. A stdio server that indexes your notes or queries a database you already have credentials for is a subprocess and a config entry: no endpoint, no hostname, no auth, no deploy. The moment the answer is a second person, you are not writing a tool any more. You are running a service.

Question two: whose credentials does it use

This is the one that decides it, and it is not a technical question. A local server inherits the ambient identity of whoever started it: their SSH keys, their cloud CLI session, their environment variables, their home directory. The permission model is that the user is the process, and it does not transfer.

A remote server has no ambient user. The docs describe Streamable HTTP as supporting standard HTTP authentication methods including bearer tokens, API keys and custom headers, and note that MCP recommends using OAuth to obtain authentication tokens. Picking one is easy. The hard part is the authorisation model underneath: which records this caller may read, which actions they may take, and how you prove it on every request.

The spec also requires that "Servers MUST validate the Origin header on all incoming connections to prevent DNS rebinding attacks", and advises binding only to 127.0.0.1 rather than all network interfaces when running locally. Local does not mean private by default: an HTTP server you meant to run only on your own laptop is reachable from the browser tabs on it.

Question three: what happens when it is down

For a local server, down means one person restarts their client. The spec's shutdown story is that the client should close the server's stdin and wait for it to exit. That is the entire incident.

For a remote server, down means everyone is down at once, mid tool call. Because protocol-level sessions and Last-Event-ID resumability are both gone, a redeploy that drops a stream is not something the protocol will reconnect for you. Microsoft's curriculum is a fair proxy for the surface area: its production server track is a 13-lab path whose labs include Security and Multi-Tenancy, Deployment Strategies, and Monitoring, none of which exist until you go remote. If nobody on your team will own that pager, you are not ready to build the remote version.

The comparison, briefly

Local (stdio) Remote (Streamable HTTP)
How it starts The client launches it as a subprocess You deploy and run it as an independent process
Typically serves A single MCP client Many MCP clients
Identity The user who started the process Whoever the token says, on every request
Data reach That machine's filesystem and credentials Whatever you explicitly grant it
Failure blast radius One person, one restart Everyone, at once
Ongoing cost Roughly none An operational owner

The migration trap

The usual plan is to build local first and make it remote later. That is often the right sequence, but the second step is rarely small: a local server has one implicit user baked into it everywhere. Paths are personal, credentials are read once from the environment at process start, and no function signature carries a caller. Going multi-tenant means threading an identity through every call site and re-deciding, per tool, what that identity may do.

State is the same story. On stdio the process was your session, so module-level variables held between calls. The spec closes that door: "MCP has no protocol-level session, so a server cannot rely on implicit per-connection state to relate one tool call to the next." Its recommended fix, an explicit handle returned by a creation tool and passed back on later calls, is a signature change to every tool touching shared state.

If you already know a team will want this, write the caller identity into the tool signatures on day one, while it still runs as a subprocess for you alone. That single decision is most of the migration.

Where this fits if you use Bespoke Prompting

Bespoke Prompting's MCP Tool build type treats this as a classification its own engine makes for you rather than a preference you state. That classification stage is deterministic, and one contradiction it flags by name is multi_user_implied: describe a stdio deployment alongside a multi-user scenario and it says so, pointing you at Streamable HTTP. DEPLOYMENT_CONFIG is one of the eight frozen sections of every MCP tool spec it produces, so the transport question is answered on paper before anyone writes the server.

Sources

Suggested internal links