#5 of 10,000+
Lakera's Agent Breaker leaderboard
The independent record of how AI agents fail
Every failure we find, across security and reliability, becomes part of one independent record that compounds with every run.
You have only ever seen your own agents break. We have seen everyone’s.
#5 of 10,000+
Lakera's Agent Breaker leaderboard
94% attack success rate
AgentHarm, official test split
Continuous vulnerability research
The corpus compounds every run
From our engagements
Every case below is a confirmed finding from an authorized engagement, reported through official channels. Targets anonymized to their vertical, failure patterns only.
Read the research →The agent’s tool connector shipped its author-trust filter off by default, so text from a stranger’s pull request and issue comments reached the agent as trusted instructions. Behind every tool call sat one full-scope token with no per-action check, meaning injected text could drive merges, approvals and secret reads with no human in the loop.
The platform’s only authorization was a caller-supplied user-id header the backend never verified. A canary account owning no data listed more than two thousand other tenants’ live support conversations, pulled their agents’ system prompts word for word, and could create and delete customer records anonymously.
Asked for analytics, the commerce agent minted a pre-signed download link to a CSV of customer records and streamed it back in a hidden tool result. The link carried no cookie, token or session. A plain request returned names, emails, locations and order totals.
Agents are failing in public
Three of the ones that got written up. Most never do.
meta ai
guessable prompt ids let one user read another user’s chats
the server never checked who owned the prompt id
cursor
a poisoned workspace turned allowlisted git commands into arbitrary execution
the allowlist was trusted, the environment it ran in was not
mexican government agencies
one operator used a consumer ai to find the holes and pull the data
195 million identities out of tax, voter and civil registries
How agents fail
Both live in one dataset
How a run works
Point us at the agent. An endpoint, an SDK, or the live channel your users already talk to.
A fleet of agents probes it in parallel. Every refusal tells the next probe where the guard actually sits.
Nothing counts until the target says it. Every finding carries the transcript that produced it.
The finding writes back to the corpus. The next run, on any agent, starts from what this one learned.
For teams shipping AI into enterprises
We find it first and write it up in the language their reviewer already uses, so the deal stops waiting on it.
For teams building and deploying agents
Drop us into the pipeline and the failure knowledge arrives as part of the build, in the same place your tests already report.
Start without talking to anyone
Running against your own production agent is scoped with us first.