APIs, integration & security — in depth

Scoping MCP Tool Permissions for Email Sending Agents

Divide email agent permissions into read, compose, and send tiers.

Senior Correspondent, Agentic Systems · · 10 min read
Cover illustration for “Scoping MCP Tool Permissions for Email Sending Agents”
Email MCP Tooling · October 7, 2026 · 10 min read · 2,350 words

The Model Context Protocol was built flexible, not safe by default, and that choice turns into a real liability the moment an agent can send email. The NSA's May 2026 guidance on MCP security design stated that the protocol was released with a design much like early web protocols, open enough to let implementers build what they wanted, but short on the specifics that would tell them how to do it safely. That tradeoff is tolerable for a tool that reads a calendar or queries a database. It is not tolerable for a tool that dispatches mail, because a send cannot be called back once it leaves the server.

MCP hands the LLM runtime ambient authority across chains of trust that span multiple hops, hops that a firewall or an identity provider was never built to watch individually. A perimeter control checks who is allowed in. An identity control checks who someone claims to be. Neither one tracks what an agent does with the authority it was granted three steps into a chain of tool calls, which is exactly where email risk concentrates. A web request that fires twice because of a retry is an annoyance somebody notices and fixes. An email that fires twice, or reaches a list of people who already unsubscribed, is a deliverability event with blocklists and spam-folder consequences, a compliance exposure under laws that govern unsolicited commercial mail, and in some cases a legal liability that lands on whoever owns the sending domain. Every section that follows is a response to that one fact: MCP does not stop an over-permissioned agent from sending, so somebody building the agent has to.

How MCP tool discovery widens the blast radius

MCP agents do not work from a fixed list of tools decided at build time. They ask the server what it can do, and the server answers at the moment the question is asked. WorkOS's 2026 overview of the protocol describes this cleanly: the server exposes its capabilities, and the agent learns about them at runtime through a tools/list request, not from anything baked into the application beforehand. That single architectural choice means the set of actions available to an agent can grow after the agent has already been deployed, reviewed, and trusted.

Picture a remote MCP server that, in its first version, exposes a send_transactional tool for password resets and receipts. A later update adds send_bulk. Every agent already connected to that server now has access to bulk sending the next time it asks what tools exist, with no new install, no new approval screen, and no developer in the loop deciding that bulk send was an acceptable addition. Wiz Research documented a version of this problem directly: a widely installed email server MCP package shipped an update that silently copied every agent-sent email to a domain the attacker controlled. The package had passed a normal review when it was first installed, because at install time it behaved exactly as advertised. The compromise arrived in a later version, after trust had already been established and nobody was looking again.

Nothing about that case required a user to make a mistake. The exposure opened the moment the tool's description entered the model's context, which is also the moment tool poisoning becomes possible: instructions hidden inside a tool's metadata can redirect what the model decides to do before any person has acted at all, and MCP as specified does no validation of a tool's definition before it reaches that context. Applied to email, the consequence is specific and ugly. An agent built and credentialed only to read campaign statistics can wake up one day able to send to the entire list it was only ever supposed to report on, and nothing in the protocol itself stands in the way.

Why over-scoping is the default failure mode

Agents built on MCP tend to end up holding more authority than their job requires by default. The protocol gives a developer no natural line to stop at, no built-in prompt that says "this is enough." The NSA's guidance notes that a large share of MCP implementations skip authentication altogether, and that even the ones that do authenticate typically lack any enforcement of roles once a session is underway. The guidance states directly that MCP "lacks support for exchanging Role Based Access Control permissions at instantiation."

For an email agent, that gap takes a concrete shape: a token issued so an agent can read campaign statistics may carry no boundary at all against a send_email call. Nobody explicitly handed that agent send access. Nobody explicitly took it away either, and in the absence of an enforced boundary, absence of restriction functions as permission. Token passthrough compounds the problem. When an upstream token gets forwarded straight into a downstream tool call instead of being exchanged for a properly scoped one under RFC 8693's token exchange model, a credential minted for one system ends up usable against a system it was never meant to reach.

This is the precondition behind what security researchers call the confused deputy attack: a server holding broad, ambient authority gets tricked into acting on behalf of an attacker, because its permission scope was wider than the task in front of it and no per-action check existed to catch the mismatch. Wiz Research identifies this absence, missing per-action authorization, as the condition that lets anyone able to influence what the model sees redirect the server's elevated privileges without authenticating as anybody.

The three capability tiers that email tool permissions must map to

Diagram: Three Tiers of Email Agent Permission. Visualizes: Visualize a strict three-tier hierarchy of email tool permissions that every MCP email agent must map to: Read (delivery statistics, list membership, suppression lookups, bounce logs — no…

The fix starts with a taxonomy simple enough to hold in your head: read, compose, and send. Every email tool an agent touches should map to exactly one of these tiers, and no agent should hold a credential that spans tiers its task doesn't call for.

Read access covers delivery statistics, list membership status, suppression list lookups, and bounce logs, nothing that writes or dispatches anything. An agent limited to this tier cannot cause a deliverability incident no matter what a prompt injection or a poisoned tool description tries to get it to do, because the capability to cause one simply isn't present in its credential.

Compose access covers drafting messages, rendering templates, resolving variables, and staging a send, but stops short of touching SMTP. Keeping compose separate from send means a malformed template, a variable that never got filled in, an unsubscribe link that's missing, gets caught before anything reaches a recipient's inbox. Pre-send linting also belongs here structurally, since the compose step only produces a candidate message, and the send step is a distinct action with its own credential, not a continuation of the same one.

Send access is the irreversible step, the only tier that actually touches SMTP, so it should carry the narrowest credential of the three. A practical implementation scopes a send key to one sending domain, one audience segment, or one message type, transactional as distinct from broadcast, rather than issuing a general-purpose key good for anything. An agent holding write access to both transactional and broadcast streams at once carries materially more risk than one scoped to a single stream. The right question when issuing any of these credentials is the minimum capability set required for the task the agent was actually built to do.

How MCP's authorization mechanisms support progressive scoping for email

MCP's authorization model is moving in a direction that matches this three-tier approach, even if it isn't finished yet. The 2025-11-25 version of the specification added incremental scope consent along with OpenID Connect Discovery support, both steps toward granting permissions progressively over the course of a session.

In practice, progressive scoping works like this: an agent starts out holding a minimal set of scopes, and if it attempts something its current token doesn't cover, the server answers with a 403 Forbidden response and a WWW-Authenticate header carrying error="insufficient_scope" along with the scopes actually required. This pattern, known as step-up authorization, is a meaningful improvement over either silently failing or silently letting the action through, because it forces a visible decision point instead of letting an under-permissioned or over-permissioned request pass unnoticed.

The shortfall for email is one of granularity. OAuth scopes as they exist today cannot always express the kind of fine-grained, tool-level permission an MCP client needs to request from an MCP server. "Compose only, no send, scoped to this one audience segment" is a precise operational requirement, and OAuth's scope model wasn't built with that level of precision in mind. Rich Authorization Requests have been discussed as a way to close that gap and allow more detailed authorization requests inside MCP, but RAR is not yet standardized in the protocol. Teams building email agents today have to make up the difference themselves, in application-layer policy. Wiz Research frames the remaining blind spot well: an identity system can tell you who made a request, but it has no way of telling you which MCP server that request passed through, what tools that server exposed at the time, or whether the tool description the model actually read at the start of its session matches what a human reviewed when the integration was first deployed.

Guardrails at the tool boundary before a send is dispatched

Because the protocol's own authorization layer stops short of enforcing anything specific to email, the enforcement has to live at the tool boundary, the last point before a request reaches SMTP, where malformed, non-compliant, or duplicate sends can still be caught and stopped.

Idempotency is the single most important guardrail in that list. An idempotency key is a unique identifier attached to a request that guarantees the operation it describes executes at most once, no matter how many times the request itself gets sent. The length of the idempotency window matters more than it might first appear. A 24-hour window handles the ordinary case, a web client retrying a failed request within seconds, but it does nothing for an agent that runs on a daily schedule, wakes up, finds no record of the previous day's run because the window already closed, and resends the same message a second time. The fix is enforcing idempotency at two separate layers at once, both at the point where the tool is called and again at the send API itself, so a gap at one layer doesn't become a duplicate send at the other.

Pre-send linting catches a different category of failure, one unrelated to permissions. Unrendered template variables, a missing unsubscribe header, an absent List-Unsubscribe link, these are the failures an agent is most likely to introduce on its own, and no amount of OAuth scoping or idempotency enforcement will catch any of them. Compliance rules belong in this same layer: one-click List-Unsubscribe under RFC 8058 is now a deliverability requirement for bulk senders reaching Gmail and Yahoo consumer inboxes, and Outlook's consumer mailboxes require a visible unsubscribe link even though they don't mandate the RFC 8058 one-click mechanism specifically. An agent that sends bulk mail without that header will run into inbox placement failures that no amount of permission scoping upstream could have prevented.

Human approval gates complete the picture alongside linting. The agent's own reasoning engine should never be the thing calling the send API directly. That interception has to happen at the tool boundary itself, enforced in code, not left to a prompt instructing the model to ask first. Risk tiering is what keeps this from becoming a bottleneck: low-risk transactional messages, receipts, password resets, confirmations, can go straight to the send API without waiting in a queue, while higher-risk outbound broadcast gets routed to a queue for review before it goes out. Purpose-built infrastructure for agent-driven sending, AgentiSend among it, has started building these patterns, idempotency enforcement, structured linting errors, tiered approval, directly into the send path rather than leaving each developer to assemble them from scratch. Review queues built this way need to move fast. A queue that takes a day to clear defeats the point of building an autonomous sender in the first place, so the gate should function as a final compliance check, not a slow editorial pass.

Per-agent credential scoping and subdomain isolation as the operational implementation

None of the tiering or the guardrails above enforces anything unless the credentials and domains underneath them are built to carry it out. Per-agent API keys and dedicated sending subdomains are how the three-tier model gets expressed in infrastructure.

Scoping needs to happen per agent, not per account. Each agent identity should carry its own key, scoped to the minimum set of capabilities its specific task requires, and a campaign-reading agent should never share a credential with a transactional-sending agent, even if both happen to belong to the same team or the same product. An email API built for this kind of use exposes structured, tool-callable endpoints with predictable JSON schemas, issues scoped keys per agent identity, returns delivery status in a form an agent can reason over programmatically, enforces idempotency keys against duplicate sends, and plugs directly into orchestration layers like MCP servers instead of requiring custom glue code to bridge the gap.

Dedicated subdomains per agent role keep reputation damage contained. If one agent's sends start generating spam complaints, that damage stays isolated to its own subdomain instead of spreading to the organization's primary sending domain or to the reputation of every other agent sending under the same name. A new subdomain still needs warmup before it carries real volume. Sending at full volume from a brand-new subdomain will trip spam filters regardless of how carefully the permissions behind it were scoped. SPF, DKIM, and DMARC all need to be set up at domain verification and kept current on each sending subdomain individually, because an authentication failure at that infrastructure layer causes inbox placement problems that no guardrail built at the application layer can repair after the fact. A platform like AgentiSend, which handles both transactional and marketing email through one API and one MCP server, removes the split-credential complexity that shows up when separate products govern separate send types, narrowing the surface where a credential gets misconfigured.

Sources

  1. Model Context Protocol (MCP): Security Design ...
  2. Specification
  3. Securing the Model Context Protocol (MCP): Risks, Controls, and Governance