Start With a Control Plane, Not a Server Blocklist

Enterprises should control Model Context Protocol (MCP) access through a centralized control plane that evaluates every tool call across four questions: who is acting, what they may do, with which data, and under what conditions. A login alone is not enough because an authenticated user can delegate authority to an AI agent whose behavior is partly determined by model output, retrieved documents, and third-party MCP servers. The practical goal is not to stop agents from working; it is to let low-risk actions proceed automatically while placing defined approval and inspection points around consequential actions. For issue-operations and case-house teams, this means an agent may be able to summarize a complaint, retrieve an approved policy, or draft a response without permission from a security administrator for every step.

Also worth reading: What AI agent governance controls should enterprises implement for AI agents in 2026? · How Should Enterprises Govern AI Agents Running Live Business Processes? · How can enterprises effectively manage the risks associated with deploying autonomous AI agents in production environments?

A useful enterprise design separates the MCP gateway or runtime from the underlying systems. The gateway should authenticate the user and workload, resolve the agent’s current permissions, inspect the requested tool and arguments, apply rate and data-loss controls, and record a tamper-resistant audit event. The target system should continue enforcing its own authorization rather than trusting the gateway blindly. This “defense in depth” approach avoids turning a newly introduced agent gateway into the single point of failure. A sensible initial objective is to inventory all MCP connections, route at least 95% of them through governed paths, and reduce unmonitored production connections to near zero within 90 days.

Enterprises should also distinguish between controlling an MCP server and controlling an individual MCP tool call. A server may expose five tools ranging from a harmless status lookup to account closure or external publication, and access to the server should not imply equal access to all five functions. Policies need to be attached to tools, arguments, data classifications, destinations, transaction sizes, and sometimes user intent. This is more precise than maintaining a list of blocked servers and is essential for support, compliance, and public-affairs platforms that must move quickly without allowing an agent to disclose protected cases or make unauthorized commitments.

Why MCP Security Differs from Conventional API Security

Traditional API security usually assumes that a developer has selected an endpoint, supplied predictable parameters, and built the business logic before the request is made. An MCP-enabled agent can select among available tools dynamically, construct arguments from natural-language instructions, and combine several calls without a predefined workflow. That flexibility turns ordinary authorization mistakes into agentic risks: prompt injection in a document may influence which tool is chosen, a broad tool may return more context than the task requires, and a sequence of individually permitted calls may produce a result that should have required higher approval.

Consider a support agent permitted to read tickets, classify cases, recommend remedies, issue refunds, and close records. It may be reasonable to allow ticket reads below a certain classification automatically, but a refund above $500 should require a different limit, while closing a regulatory case should require human confirmation. A conventional role might permit all three actions because the person’s job function requires them; an MCP policy can instead evaluate the particular transaction at runtime. Support teams can encode rules such as “draft without approval,” “review before external delivery,” or “block entirely,” which are often more useful to case managers than infrastructure-level network permissions.

The second difference is that agents are non-deterministic. A deterministic application produces the same output for the same inputs, but an LLM may choose different tools or arguments after minor changes in wording, retrieved content, or tool descriptions. Security teams therefore should not treat a successful test suite as proof of stable behavior. They need continuous evaluation using representative cases, adversarial inputs, and production monitoring. For an issue-operations platform, that could mean 100 test prompts for media inquiries, policy escalations, customer complaints, and sensitive data requests every time a model, system prompt, tool schema, or major policy changes. A target of zero unauthorized external sends and zero confirmed cross-tenant disclosures should remain non-negotiable even while efficiency improves.

Use a Layered Policy Model for Every Tool Call

A mature policy model should combine identity, purpose, tool, data, action, and session context. Identity answers whether a known employee, service account, customer, or agent is responsible. Purpose indicates whether the agent is handling a support case, conducting compliance research, preparing a public-affairs response, or performing an unrelated task. Tool and data controls then determine which system can be queried and which fields can be returned. Action controls govern whether the result is merely a draft or is sent, published, paid, deleted, or used to change a case status.

These dimensions should be evaluated in sequence rather than collapsed into a broad role. An employee may be allowed to read a case but not export it; an agent may be permitted to draft a reply but not send it; a vendor may access a limited schema but not underlying customer records. Policies can be static, such as a prohibition on modifying legal holds, or contextual, such as requiring a compliance officer’s approval whenever a response cites a restricted policy. Contextual controls are slower to design but prevent a single broad entitlement from producing unnecessary business risk.

Control layerMain questionExample enforcement for a case-house workflowTypical latency treatment
IdentityWho or what is acting?Bind the user, agent, tenant, service account, and sessionAdd seconds; cache carefully
Tool authorizationIs the requested operation allowed?Permit case search, but block automatic case closureMilliseconds
Data policyWhat may be read or returned?Mask phone numbers, health details, and legal notesOften milliseconds; heavier for live filtering
Action policyCan the agent commit the result?Draft a regulator response; require counsel to send itHuman review only for defined actions
Session and risk policyDoes the overall sequence look acceptable?Stop repeated exports or unusual tool chainsAdd seconds when behavior is unusual
A useful operating principle is to make the safe path fast. If routine case lookups take 20 seconds because every request requires a security ticket, teams will bypass the gateway or create shadow tools. Policies should be pre-approved by data owner, tenant, tool, and common action class wherever possible, with real-time evaluation reserved for exceptions. That allows roughly 80%–90% of low-risk operations to proceed automatically while a much smaller set of consequential actions receives deeper review. Exact percentages will vary, but the design should measure and publish automatic versus escalated calls rather than treating governance as a universal brake.

Put Human Approval at the Commit Point

The most effective human approval is usually placed immediately before an irreversible or externally visible action, not at the beginning of an entire agent session. A reviewer does not need to approve every search, classification, or draft, but should see the final text, intended recipient, relevant source material, and the action the agent intends to take. For external communications, the approval record should include the channel, audience, attachments, links, and any sensitive information included in the body. In a public-affairs workflow, this is the difference between allowing research automation and preventing an agent from making an unreviewed statement on behalf of the organization.

Approval interfaces should show evidence, not just ask someone to click “approve.” The reviewer should be able to inspect the tool calls that produced the recommendation, view the data sources and policy versions used, and compare the proposed communication with the underlying case record. Highlighting unsupported claims, confidential passages, regulated topics, and changes from an approved template can reduce review time. If an agent proposes publishing a response to 50,000 customers, the interface should also show recipient count, jurisdiction, channel, campaign classification, and whether legal or communications approval is required.

Timeouts and defaults matter because every exception needs an operational outcome. A policy might require approval within 15 minutes during business hours and otherwise hold the action without sending it. Silence should not be interpreted as consent, and a failed approval service should fail closed for high-risk actions. Organizations can set service levels such as approving routine external drafts within four business hours and urgent safety communications within 30 minutes. Measuring these intervals keeps governance aligned with the speed of the underlying issue operation instead of forcing every team into the same approval queue.

Treat Identity and Agent Provenance as Separate Problems

MCP access requires identity for both the human or service that initiated the work and the agent executing it. If the gateway sees only an API key shared by an entire department, it cannot determine which employee delegated a task or whether that employee still has access. The better pattern is to issue short-lived, workload-specific credentials after the user signs in, then propagate user, tenant, session, and agent identifiers through each tool call. Service-to-service credentials should also be distinct from user credentials so that support automation cannot impersonate a named employee.

A2A, the Agent2Agent protocol, addresses communication between agents, while MCP primarily connects agents to tools and data. Enterprises should not conflate the two. A2A messages may require sender authentication, task authorization, message integrity, and limits on delegation, but those controls do not automatically authorize the MCP tools an agent will later call. The permission granted in one protocol must be respected in the other. If a support agent can ask a research agent for an answer, the original user’s data restrictions should still apply to any MCP call required to produce it.

Delegation should be explicit and bounded. An agent may receive a capability such as “search approved public-policy sources for 24 hours” rather than inheriting all of the manager’s access. High-risk tools should not be silently delegated through multiple agents because the number of hops can obscure accountability. For audit purposes, the system should retain the initiating user, every intermediary agent, each approval, and the final destination. A useful minimum is to preserve identity and action metadata for at least 12 months, with legal, contractual, and regulatory requirements determining longer retention.

Discover, Register, and Continuously Audit MCP Connections

The first practical step is to find MCP servers and tools that are already operating, including those embedded in developer tools, desktop clients, browser extensions, and vendor platforms. Network traffic alone is insufficient because local or proxied connections may not resemble recognizable internet traffic. Teams should combine endpoint inventory, DNS and egress logs, source-code searches, cloud configuration records, procurement records, and interviews with product owners. MCP scanners can help identify exposed servers, but a scanner result should be validated manually before a server is classified as production-critical.

Every approved server should have a registry entry naming its owner, business purpose, tenant boundary, tools, data sources, authentication method, data classification, external dependencies, and incident contact. The entry should state whether the server is read-only, can modify records, can communicate externally, or can execute code. A support vendor whose tool only searches approved help articles should not be classified the same way as a server that can issue refunds or publish public statements. Registries reduce the chance that a useful tool remains outside procurement, security, privacy, and records-management oversight.

Continuous auditing should compare declared behavior with observed behavior. Security teams can alert on new tools appearing, permissions changing, unusual argument patterns, sensitive fields being returned, or a server moving to a new domain. For a medium-sized enterprise, a reasonable initial target is to review all new production registrations within two business days and investigate all high-severity alerts within one hour. Organizations should also re-attest every six months and immediately after a major model, tool, vendor, or data-source change. A registry that is accurate only at launch becomes a historical directory rather than a control.

Compare Gateways, Proxies, and Platform-Native Controls

An MCP gateway is typically the policy enforcement point between agents and servers, while a proxy may focus more narrowly on network routing, protocol translation, or budget enforcement. Both can contribute to governance, but they are not interchangeable labels. A gateway may provide identity-aware authorization, data filtering, approvals, logging, credential brokering, and tool discovery. A proxy may provide egress filtering or spend limits while lacking the business context needed to decide whether a particular support response is appropriate. Some enterprise platforms also offer registries and governed access controls, but buyers should verify whether those features are native, optional, or dependent on a separate vendor.

ApproachStrongest useCommon weaknessBuying or operating question
Endpoint controlsPreventing malware or unauthorized network accessLimited visibility into arguments, intent, and resulting business actionsCan it identify individual tool calls and data fields?
MCP gatewayCentral authorization, approval, filtering, and auditCan become a latency bottleneck or single failure pointWhat happens when the gateway is unavailable?
API-native authorizationProtecting each underlying systemTeams must map agent behavior to many separate API policiesAre policies centralized or reimplemented per service?
Platform-native controlsConvenient governance inside one AI platformMay create vendor lock-in or uneven coverageCan controls be enforced across other models and servers?
Local agent permissionsFast development and narrow experimentationOften invisible, inconsistent, and difficult to revokeHow are local credentials discovered and brought under policy?
No single product category eliminates the need for system-level authorization. MCP controls should narrow what an agent can attempt, while the case system, database, ticketing platform, or publishing service remains responsible for validating the final operation. Buyers should test fail-closed behavior, tenant isolation, log export, policy simulation, credential rotation, and support for at least 10–20 common governance conditions. They should also calculate added latency at realistic concurrency; a control that adds 30 milliseconds at low volume may add several seconds when a gateway serializes thousands of agent operations. The correct comparison is not feature count, but whether the platform can enforce the organization’s actual risk decisions without making routine work unacceptably slow.

Avoid the Mistakes That Turn Governance Into an Obstacle

The most common mistake is treating MCP adoption as a shadow-IT problem and responding with a blanket prohibition. Developers then create personal servers, route traffic through unapproved providers, or embed sensitive data in prompts and tool descriptions. Another common error is assuming that model safety policies provide the same protection as enterprise authorization. A model may be instructed not to expose secrets, but it can still be manipulated by retrieved content, and it does not replace server-side permission checks or transaction approvals.

Organizations also make the mistake of granting agents standing human privileges. If an agent uses a shared administrator credential, there is no reliable way to distinguish an authorized workflow from an induced action. Broad read access is similarly dangerous because a response can disclose more than the user asked for. Data minimization should be built into tools: return case IDs and status by default, provide contact details only when necessary, and redact sensitive fields before model processing whenever possible.

The final mistake is assuming that human approval is effective without designing the review experience. Reviewers who receive hundreds of generic alerts will approve mechanically or reject everything. Policies should therefore be precise, exceptions should be measurable, and approval requests should include the exact output and proposed action. For example, a rule that flags every use of the word “confidential” will create noise, while a rule that checks attachments, data classification, recipient scope, and source access can identify a real risk. Good governance reduces the number of decisions people must make; it does not transfer the entire judgment to a gateway.

Establish a 90-Day Rollout and Know When to Pause

Enterprises do not need to perfect an agent-governance program before allowing controlled use. During the first 30 days, they can inventory known MCP connections, nominate owners, identify systems containing regulated or confidential data, and route development environments through a basic gateway. By day 60, they should have a registry, short-lived agent identities, default-deny rules for production writes, data filtering, and centralized audit logs. By day 90, the objective should be to govern at least 95% of production tool calls, test high-risk workflows, set approval service levels, and document an emergency shutdown procedure.

Read-only workflows should usually move first because they can be monitored and corrected without changing external state. Drafting support responses or compiling public-affairs research is also a reasonable starting point if retrieved material is filtered and no message is sent automatically. More dangerous actions—payments, deletions, case closure, account suspension, regulatory submissions, and public posting—should remain disabled or require explicit approval until teams have measured reliability. A useful launch gate is 30 days of stable operation, fewer than 1% of calls requiring manual repair, and zero confirmed unauthorized disclosures or external commitments.

There are circumstances in which an organization should stop and redesign rather than wait for a perfect policy engine. It should pause if the gateway cannot preserve tenant boundaries, if audit records can be altered without detection, or if agents need permanent administrator credentials to function. The same applies when a tool exposes an entire datastore to a model but cannot support field-level filtering, or when approval timeouts default to approval. Security leaders should also intervene if the business cannot identify the server owner, the source of returned data, or the human responsible for an external action.

The broader operating principle is controlled autonomy: automate the reversible and routine, review the consequential, and prohibit the unaccountable. Applied to MCP, that approach allows issue-operations and case-house teams to move faster because routine analysis and drafting happen without repeated security tickets. It also gives compliance and public-affairs leaders evidence that access was limited, decisions were traceable, and humans remained responsible at the point where authority becomes real.