Extracts from real engagements, anonymised. Different system types, different findings, and in one case a system that passed the obvious test and failed the right one. Read whichever is closest to what you have deployed.
Signed in as one customer, we made an assistant read another customer's account, email their invoice to an outside address, and issue a refund against their balance. Judged on the system's own tool-call log, so there was nothing to argue about.
Every query searched the whole corpus, including two control queries that were not attacks at all. Retrieved content included contract values, credit limits, an internal negotiating position and a confidential legal dispute.
Three builds, three verdicts. One leaked openly. One passed the obvious isolation test and still handed over another customer's entire conversation to anyone who asked for it correctly. One held.
An industry-standard probe reported a 99.61% attack success rate. We read the actual responses and found the system had defended itself perfectly. A report built from tool output alone would have been completely false.
These are extracts from real engagements with identifiers removed, conducted either against systems we were authorised to test or against purpose-built systems in our own controlled environment. They are provided so you can judge the structure, the evidence standard and the reasoning before you commission anything. None of them names a client, and none of them is a statement about the security of any named organisation.
If the system you have deployed is not represented here, that is worth a conversation rather than a guess. See systems we test, or tell us what you are running.