Agents Broke Your Access Model and Nobody Filed a Ticket
Every access-control system your company runs was built on three assumptions: there is a user at the front of every request, the thing making the call is that user, and the scope of the call is what the user consented to.
An agent calling a sub-agent calling an MCP server breaks all three at once, and it breaks them quietly. Nothing throws an exception. The request looks valid at every layer. You find out during an audit, or during an incident.
The agentic call path
Jane, a support engineer, types one request into her support copilot: “Prepare a refund summary for ticket #4521.”
The copilot’s planner agent delegates to a research agent, which pulls the ticket from Zendesk and the linked bug from Jira. It also delegates to a billing agent, which fetches the payment record from Stripe. One human request has become six machine-to-machine calls. Each agent and each MCP server may be owned by a different team or a different vendor. Only the first hop had a human behind it.
Call the chain user → application → agent → sub-agent → … → MCP tool the agentic call path. It’s the unit that has to be governed, and by the third hop three questions have no clear answer:
- Who is this call for? Jane authenticated three hops ago. Is her identity still on the request, or has it been replaced by a service account?
- Who is making it? A bearer token proves possession of a credential. It does not prove which of your two hundred agents is holding it.
- What may it do? A token minted for the first callee is over-privileged for every callee after it.
Answering all three at every hop, with an audit trail, is what agent governance means. Everything else is detail.
Four failures you will hit
Shadow agents. Teams build on whatever they like: LangGraph, CrewAI, Bedrock, a homegrown HTTP service, a copilot inside a SaaS product. Nothing forces them to declare it anywhere. A growth team ships a CrewAI agent that queries Salesforce with a shared API key; it works, and no central system knows it exists. It can’t be found, audited, rate-limited, or switched off. Even a copilot tucked inside a SaaS product can carry capabilities nobody on the team knows about, the way the hidden slash commands in Notion AI surprise people who only use the obvious buttons.
Unattributable actions. When an agent forwards the user’s token, downstream logs show only “Jane.” When it uses a shared service account, they show only “the service.” After an incident, “the retrieval agent exposed that data” is unprovable either way.
Chained authority nobody approved. Three agents in a procurement flow: one creates vendor records, one retrieves supplier banking data, one initiates payments. Each is defensible alone. Chained, they form an end-to-end payment path no reviewer would have signed off on. This is separation of duties re-broken by software, and no single system sees the whole chain because each only sees inside its own boundary.
The confused deputy, at scale. An agent with broad access takes instructions from content it read. Prompt injection in a Jira ticket becomes a real Salesforce write, because the agent’s token doesn’t distinguish “what Jane asked for” from “what the ticket text said.”
Agent identity is a third kind of principal
The root mistake is modeling agents as something they aren’t. Teams reach for one of two existing shapes:
The agent acts as the user (forwarding their token). Now the agent has everything the user has, attribution is destroyed, and a single injection gets the attacker the user’s full reach.
The agent acts as a service (shared account). Now every user of the agent gets the union of all permissions the account holds, rotation means touching every agent, and logs identify a service rather than an actor.
An agent is neither. It’s a third kind of principal:
| User | Service account | Agent identity | |
|---|---|---|---|
| Represents | A person | A fixed service | A registered agent |
| Acts as | Itself | Itself | Itself, or on behalf of a user |
| Behavior | Human judgment | Fixed configuration | Autonomous (picks tools, chains calls) |
| Governance need | SSO, RBAC | Rotation, inventory | That, plus delegation rules, ownership, per-hop attribution, kill switch |
Giving each agent its own verifiable identity makes the rest possible:
Attribution. The receiver can tell whether the caller is a human, a service, or which agent. Actions become provable.
Per-agent policy. One user drives many agents, and they shouldn’t all inherit that user’s full reach. Jane’s support copilot reading Jira and her engineering agent writing to it are different principals even though both act for Jane.
No anonymous agents. If an agent can only reach a tool by presenting a registered identity, registration becomes the enforcement point. You get an org-wide inventory as a side effect, and you can revoke one rogue agent without touching the rest. Think of Chrome’s Task Manager: one keyboard shortcut lists everything running, including the tabs you forgot were open.
Registration should carry more than a name: an accountable owner, a one-sentence business purpose, a defined authority (which tools and data domains, which actions, which users it may act for), and a review cadence. Defined authority is the reference point everything else gets measured against. Without a declared boundary there’s nothing for runtime behavior to drift from.
Delegation, not impersonation
The mechanism that makes per-hop scope work is token exchange. Instead of forwarding the inbound token, each hop presents its own identity plus proof of the delegation it received, and receives a new token scoped to that callee alone, for that user alone, valid for that hop only.
The result is a request that says: “agent support-copilot, acting for user jane@, delegated from helpdesk-app, may read Zendesk tickets, expires in 60 seconds.” That token is useless to any other callee and useless a minute later. Expiry by default is catching on in consumer software too, and the automatic reboot security feature on Android is a small example of a system quietly returning itself to a safer state.
Standards here have matured faster than most people realize. OAuth token exchange has been available for years. Cross App Access (ID-JAG) addresses the agent-to-app case. Microsoft Entra Agent ID and Okta’s agent features issue first-class agent identities. SPIFFE/SPIRE covers workload attestation where you want the runtime itself proven. The pieces exist; the integration work is what’s missing.
Where enforcement belongs
You cannot enforce this in the agents. Agents are the thing being governed, they’re written by many teams on many runtimes, and asking each one to implement delegation correctly guarantees the weakest implementation defines your posture.
Enforcement belongs at a proxy every hop traverses, called an agent gateway. It validates the caller’s identity, checks the delegation against the agent’s registered authority, mints the scoped downstream token, applies guardrails to the payload, and records the hop. Because enforcement and logging happen at the same point, the audit trail is a byproduct rather than a separate project, which saves your team a lot of time.
TrueFoundry’s writeup on why agent governance breaks the enterprise access model is the clearest statement of the problem I’ve found, and their key concepts page is worth reading even if you build your own gateway. It separates the identity provider, authorization server, and enforcement point roles cleanly, which is the distinction most implementations get wrong. Conflating them is why teams end up with an IdP that issues agent tokens but no per-hop scoping.
A staged plan
Inventory first. You cannot govern what you can’t enumerate. Registration gated at the tool-access layer gets you this without asking teams to volunteer.
Then identity. One verifiable identity per registered agent, distinct from users and service accounts.
Then per-hop scope. Token exchange at each hop. Start with your highest-risk destination: whichever MCP server touches money or PII.
Then guardrails and audit. Inspect payloads for injection, PII, and unsafe tool calls at each hop; log every hop including denials. Knowing how prompts steer a model, as covered in this look at practices for training AI models with prompts, makes it easier to decide what those guardrails should catch.
The order matters because each stage makes the next one cheaper, and because the first one is achievable this quarter. Most teams can’t currently answer “how many agents do we have in production.” That’s the thing to fix before anything else, and it saves the most time later.








