Get a demo
INTO THE EXPANSE · chapter 04

The Agent Is not the Target. It’s the Way In.

To an attacker, the agent is a path to everything it can reach.

Chapter 04 episode preview
Coming Soon

The economics of intrusion

Ask the market what it means to secure an agent and the answers point to the agent itself: scan its model, harden its system prompt, filter its inputs, guardrail its outputs. All of this is useful. None of it is the point.

Consider how an attacker sees the same agent. Not as a chatbot to break, but as something better than a stolen credential: a pre-authorized identity with standing access to email, files, records, and pipelines. It's already trusted and already connected. It works at machine speed. And it runs on a model that changes with every vendor release and carries security flaws of its own. To an attacker, the agent is a path to everything it can reach.

This changes the economics of intrusion. A traditional attacker has to fight through every stage of a kill chain — get in, escalate, move laterally, find the data, get it out — and every stage is a chance to catch them. A compromised agent skips most of that. Access is pre-granted, the environment is pre-mapped, and the data movement looks like the agent doing its job. The GTG-1002 campaign is the concrete version: attackers manipulated an AI system into autonomously executing 80–90% of an espionage operation across thirty targets, at thousands of requests per second. Security tooling built for a different era watched it happen.

The most accurate frame for agent risk turns out to be an old one: insider risk. Irregular's Dan Lahav calls AI "the frontier of insider risk,"1 and Reco's research came to the same conclusion from the enterprise side — an agent operates with the trusted internal access once reserved for employees and service accounts, without the background check, the manager, or the sense of self-preservation. Securing an insider was never about monitoring what they say. It's about knowing what they can access, watching what they do with it, and noticing when the pattern changes. The same holds for agents: an agent’s risk is its reach.

An agent’s risk is its reach.

Prompt injection: the honest version

No chapter on agent security is credible without saying this plainly: prompt injection is unsolved, and at the input layer it is probably unsolvable.

The reason is structural. An agent is useful precisely because it ingests untrusted input and treats it as material to act on. That input arrives two ways, and both matter. Direct exposure is anyone who can send data straight into the agent's instruction channel — who can chat with it, what events trigger it, which other agents can invoke it. Indirect exposure is everything the agent reads while working: web pages, files, cases, connected sources, any of which an attacker can poison in advance. Direct exposure is classic prompt injection, where the attacker talks to the agent. Indirect exposure is quieter. The attacker never touches the agent at all, just leaves instructions where the agent will find them.

The OWASP Agentic Skills Top 10, the industry's emerging consensus framework, classifies untrusted external instructions as AST05.2 Reco researchers Daniel Alfasi and Tal Shapira are listed contributors, and the attack scenario Reco contributed from its own red teaming work shows how chains make the problem worse. They call it relay-node amplification: an instruction injected into one agent gets amplified as it moves down a chain of agents, because downstream agents extend trust to upstream ones. The injection doesn't have to defeat every agent in the chain. It has to defeat the least defended one.

Security researcher Simon Willison gave the worst case a name: the lethal trifecta.3 An agent with access to private data, exposure to untrusted input, and a channel to send data out. Any two of the three are manageable. All three together mean an injection doesn't just land — it exits with the goods. When the same agent can also write, execute, or administer, the trifecta stops being a leak and becomes an attack platform.

An agent with access to private data, exposure to untrusted input, and a channel to send data out. Any two of the three are manageable. All three together mean an injection doesn't just land — it exits with the goods.

Filters catch the crude attempts, but they won't catch them all, because distinguishing "content to summarize" from "instructions to follow" is exactly the judgment agents get wrong under adversarial pressure. And after relay-node amplification, the malicious instruction may not even come from outside. It may arrive from a trusted peer agent that was compromised one hop upstream.

If some injection will always get through, the useful question is not how to stop every malicious prompt. It's what happens when one lands.

Run the Agentforce chain from Chapter 2 (public internet → email → case record → agent prompt) against two different organizations. In the first, the case agent can read and write cases. An injected instruction lands, and the blast radius is a support queue. Embarrassing, but contained. In the second, the same agent also holds a connection to the data warehouse and a knowledge-base integration, because those made it more helpful. The identical injection now reaches customer records and publishes to systems humans read and trust. Same attack, same platform, same prompt filter. The difference between an incident report and a breach disclosure was decided months earlier, by permissions and connections nobody was tracking.

So the honest posture on injection is: assume it, and make it survivable. Survivability comes from the unglamorous controls discussed in Chapter 3 — least privilege actually enforced, connections actually mapped, the blast radius known before the attacker traces it for you.

Toxic combinations

The industry analyzes agents one at a time. Attackers work in chains.

Reco's term for the pattern is a toxic combination: agents and access rights that are individually reasonable and collectively an incident waiting to happen. The canonical example involves three tools nobody would flag alone. Say, an Agentforce agent built by RevOps writes Zoom call summaries into a Notion database, which is sensible. A customer-success copilot publishes that database to a public knowledge base, which is also sensible — the database was for shareable content. Together, they publish internal sales conversations, customer names, and deal terms to the open internet, with no injection, no compromise, and no misbehavior by either agent. Two safe agents composed into an exposure, and nobody was looking at the level where the composition happened.

Slack Code produces these chains as a feature. Conversation → code channel → agent → GitHub → production deploy → preview URL: five trust boundaries crossed in one workflow, each hop individually approved by someone, the full path approved by no one. Agent-to-agent interaction is reshaping how work moves through the enterprise, which means chains now form between agents directly, without a human hop where a review could anchor. To an attacker, every agent-to-agent trust link is a potential amplifier. The worst configuration — an agent that other agents can trigger, that holds privileges, and that can act on other agents — is the precondition for worm-style lateral movement.

Chains also break the logging that incident response depends on. A user asks a low-privilege agent for something harmless. That agent asks a second agent with broader access. The second agent performs the sensitive action. The logs show that the second agent had permission, and nothing about who originally asked or why it acted. Authority can be preserved through a delegation chain, or expanded, or quietly laundered — and every hop is a place where a limited request can pick up privileges its originator never held. Evaluating the last hop tells you the action was authorized. Only the whole chain tells you whether it should have been.

This is why agent security can't be accomplished agent by agent. A per-agent review of either agent above approves it. Seeing that two rights combine to form an exposure requires a view that holds the whole landscape — every agent, every permission, every connection — and keeps it current.

Risk that explains itself

All of the above raises an operational question: with thousands, or millions, of agents and a security team that hasn't grown, what do you look at first? Hand-reviewing everything is impossible. Risk scoring is the answer.

Four principles separate scoring that works from scoring that doesn't.

  1. Risk is a conjunction, not a checklist. An agent's risk is what it could do multiplied by whether anyone can make it do that. A powerful agent nobody can reach is dormant. A wide-open agent that can touch nothing is noise. Only the conjunction is dangerous: capability meeting reachability meeting an asset worth taking. Scoring that adds up findings will rank ten trivial flags above one complete attack path. Scoring that multiplies exposure by reach gets the order right, but it requires the entire ecosystem view, because "can anyone reach it" is not computable from the agent's own settings.

  2. The same finding is not the same risk. Take one concrete finding: credentials discovered in an agent's system prompt. On a dev-only sandboxed agent with output guardrails, that's a medium-severity cleanup item. On a customer-facing agent with no guardrails, running on a jailbreakable model, the identical finding is critical. Severity that ignores the agent's context doesn't match reality, and a queue of context-free findings is sorted wrong.

  3. Critical should mean a chain exists. When an agent lands in the top tier, the reason should be a named, complete, reachable attack path that an analyst can read and verify (i.e., this capability, reachable this way, touching this asset). An aggregate score can accumulate its way to "critical" without any actual path existing, which is how alert fatigue gets rebranded as prioritization. There's a simple test to put to any risk model: for this agent, show me the chain. If the answer is a heat map, the model is describing anxiety, not risk.

For this agent, show me the chain. If the answer is a heat map, the model is describing anxiety, not risk.
  1. Controls attenuate risk; they never erase it. Human-in-the-loop approval, input validation, structured inputs, hardened prompts — all real, all worth deploying, all worth crediting in a score, since a score should reflect residual risk rather than raw exposure. None of them is a reason to stop watching.

Intent, done at the right level

The market has begun talking about agent intent — inferring what an agent was built to do, and flagging drift when its behavior diverges. The idea is right, and it deserves more than buzzword treatment, because done properly it's how you catch an attack early instead of reading about it in the postmortem.

Start with what intent is not: what the prompt said. A prompt is a claim. Intent is a judgment about the purpose behind an agent's requested action, made against who the agent is, what it's allowed to do, what it's trying to touch, and what its legitimate job looks like. The same mechanical action reads differently depending on that context. An HR onboarding agent creating an employee ticket is a workflow executing; the same agent querying sales commission data is scope drift. A support agent summarizing a customer case is an average Tuesday; the same agent exporting the customer base to an external workspace is exfiltration wearing a work badge. A finance agent updating an invoice is expected; a finance agent minting itself a new privileged OAuth connection may be the first visible move of a compromise.

In every one of those pairs, the prompt could look impeccable. The judgment comes from context: the agent's identity and owner, its permissions, its normal behavior, the sensitivity of the data, the destination, the policy that defines its role. Unwanted intent is drift from the agent's approved role, delegated authority, and data scope — even when the prompt itself looks benign. That last point matters most, because injection-driven attacks are designed to look benign at the prompt layer. They only become visible as intent violations at the ecosystem layer.

This is also why intent without reach is a horoscope. Knowing a calendar agent "intends" to manage meetings tells you nothing actionable until you know it can also touch the HR system — at which point the gap between intent and reach is the finding. Agent-to-agent systems raise the difficulty again, because intent becomes transitive. The question stops being "is this agent allowed to do this?" and becomes "was it legitimately asked, by an actor whose own authority covers the request, on behalf of someone whose authority covers it too?" Answering that means holding the whole delegation chain. Context is what makes intent a control instead of a caption, and the graph is what makes the context computable.

Red teaming and scanning face the same requirement. Generic jailbreak batteries measure a model's manners. Meaningful adversarial testing attacks the agent in its ecosystem — this org's intake features, this agent's actual scopes, these adjacent agents. And the model itself has to be part of the test. Reco's red teaming work across 39 model backbones found that identical agents fail differently depending on the model underneath, a finding contributed to the OWASP framework as AST08, model-dependent injection resistance. Failure is configuration-specific. You have to test the agent you actually deployed, on the model it actually runs, in the environment it actually inhabits.

What securing an agent actually means

Run the IPCA framework at full depth and agent security becomes concrete.

The IPCA framework applied to agent security: the question each pillar asks and the control that answers it
IPCA pillarThe questionThe control
IdentityIs this agent registered, owned, and distinguishable from every other actor?An identity per agent (not a shared service account) with a named owner and a lifecycle: created, reviewed, retired when its owner leaves or its purpose ends.
PermissionsDoes its access match its function still?Least privilege enforced against observed need, re-checked on a cadence, because scopes granted in March could be the blast radius in August.
ConnectivityWhat can it reach, transitively, and what chains does it participate in?The full graph — OAuth grants, integrations, agent-to-agent links — with toxic combinations detected as they form.
ActivityIs what it's doing normal for what it is?A behavioral baseline per agent, judged in context: its intent evaluated against its role, its reach, and its delegation chain.

Of the four, Activity is the one that resists a clean answer. Identity, Permissions, and Connectivity can mostly be reconstructed from configuration: who an agent is, what it's allowed to do, what it's wired to. Activity means watching what actually happened, and an agent's real behavior rarely stays inside the platform it originated from. A Copilot agent reaching Salesforce through an MCP connector doesn't generate Salesforce-side telemetry of its own accord. Its actual behavior, the records touched, the fields changed, the data exported, shows up as events inside Salesforce, which Microsoft's own logs never see. Building a real behavioral baseline means correlating telemetry across every platform an agent's identity touches, not just the one it started in. It's the same lesson Chapter 1 and Chapter 3 already established for discovery: no single vantage point sees the whole thing.

Machine speed cuts both ways

One more property separates agent security from what the enterprise has secured before: failures correlate.

Human-era incidents were mostly independent events. Agent-era failures share platforms, packages, models, and connectors. When LiteLLM's release pipeline was compromised for forty minutes, CloudSEK's analysis put the blast radius at 434,000 CI/CD pipelines and over 2,500 organizations.4 2,500 organizations inherited the breach at once — not because 2,500 security teams made 2,500 mistakes, but because one shared component failed and the ecosystem transmitted it. Lahav's framing for the coming period is severity, scale, and simultaneity. The agent ecosystem is the transmission medium for all three: the same OAuth patterns, the same MCP servers, the same handful of model providers, the same platform triggers, deployed across every enterprise at once. A technique that works against one Agentforce configuration works against thousands.

That's the case for the architecture this series has been describing. A threat that propagates through shared ecosystems has to be met by defense that sees the ecosystem, continuously and in context. Per-agent inspection isn't wrong. It's just aimed at the smallest part of the problem.

What none of this does is act in the moment — catch a risky action mid-flight and stop it before the data moves. That's runtime, the loudest promise in the market and the layer with the most structural blind spots. One of them now carries an OWASP designation: AST09, the unreachable skill, an agent capability deployed inside a third-party platform where no external scanner can go. Chapter 5, we’ll discuss what runtime can do, what it can't, and why the gateway you deploy to watch your agents might be the way in.

Download this chapter

Take Chapter 04 with you

Thank you! Your submission has been received!
Open the PDF
Oops! Something went wrong while submitting the form.

Get Started

Be the team that enabled agents without losing control.

Thank you! Your demo request has been received.

Prefer to look around first?Take the product tour →

Take Chapter 04 with you.

Get the chapter as a PDF, or talk to the team.