Provok Book a scoping call
AI red teaming, independent of the builder

Break your AI
before someone
else does.

Adversarial testing for the chatbots, copilots and agents you've put in front of customers and staff. We play the attacker, document every crack, and nothing leaves Australia.

Signed authorisation
Staging by default
Insured for offensive testing
Onshore, nothing leaves Australia
SESSION // PROMPT_INJECTION ● CRITICAL
ILLUSTRATIVE EXCHANGE · NOT A REAL CLIENT SYSTEM
The problem

Your team built it.
Your team can't test it.

Even a strong internal team can't independently attack the system it just shipped. That isn't a competence gap, it's how assurance works. The people who know where the guardrails are aren't the people who'll find the ways around them. An outside adversary will. Better that it's us, on your authority, than the internet on its own.

What we test

The attack surface your pen test skipped.

Traditional testing checks the network and the API. It doesn't check what the model can be talked into. We test the AI layer itself, mapped to the OWASP Top 10 for LLM applications.

LLM01
Prompt injection
Direct and indirect. Hidden instructions in documents, tickets, web content and tool output.
LLM07
Guardrail bypass
Jailbreaks that push the model past the limits you thought you'd set.
LLM06
System prompt leak
Extracting instructions, keys and internal logic the model was never meant to reveal.
LLM02
Data disclosure
Coaxing out training data, other users' data, and sensitive records in scope.
LLM08
Tool & function abuse
Turning the agent's own tools and integrations against the business.
LLM09
Excessive agency
Actions the agent can take that it should never have been allowed to.
How an engagement runs

Scope. Attack. Report. Retest.

Fixed scope, fixed price, delivered in days. One team plays the adversary, a separate analyst documents every finding. The attacker never writes its own report. See a full walkthrough →

01

Scope and authorise

We agree exactly what's in scope and sign clear rules of engagement before anything is touched. Staging by default, production only on your explicit written authority.

02

Adversarial testing

Automated attack tooling plus analyst-directed techniques, run against your AI in an isolated, fully logged environment. Every attempt captured as evidence.

03

Independent report

Findings, severity, and clear remediation direction your engineers can act on. Mapped to the OWASP LLM Top 10 and the governance framework you already answer to.

04

Retest

Once you've fixed what we found, we test again and confirm the fixes actually hold. Included, not an upsell.

Onshore by design

Your systems are sensitive. Your data has rules.

Testing runs on Australian infrastructure. Evidence is held and destroyed in Australia. The work is mapped to the frameworks your auditors and enterprise buyers already ask about, so a report you can hand straight to them, not translate first.

AU Voluntary AI Safety Standard
AU Guidance for AI Adoption
ISO/IEC 42001
Privacy Act 1988
OWASP LLM Top 10
NIST AI RMF
How we work

Built to survive an audit.

Independent
Provok does not test systems built by Evolaition, our AI automation service line. That work goes to an external party. An assessment of your own work is worth nothing, so we do not produce one.
Authorised
Signed authorisation and rules of engagement on every job. Nothing touched without it.
Onshore
Evidence handled and destroyed in Australia, within Australian jurisdiction.
Insured
Professional indemnity and cyber liability covering offensive testing.
Who it's for

If your AI failing would actually matter.

Any company running AI that takes real user input. If a breach, a leak, a wrong answer or an action taken on someone's behalf would cost you more than a bad review, you're who we test for. Regulated or not.

Chatbots. Copilots. Agents. RAG.

Extra weight where the stakes are highest, government, health, finance and other regulated sectors, where AI-specific assurance is fast becoming something auditors, boards and enterprise buyers expect.

Why us

Anyone can run a scanner.

The tools are free and public. What they produce is noise until someone with judgement reads it. That reading is the work, and it is what you are actually paying for.

From a real engagement 99.61%

The scanner said the system was compromised. It wasn't.

An industry-standard jailbreak probe reported a 99.61% attack success rate across 1,280 attempts. We pulled the actual responses and read them. Not one showed the system breaking role. It had defended itself perfectly and been scored as a total failure. A report generated straight from tool output would have told that organisation's board something completely false.

From a real engagement 12 / 14

Against an AI with tools, we made it move money.

Authenticated as one customer, we made an assistant read another customer's account, email their invoice to an outside address, and issue a refund against their balance. Twelve of fourteen attacks succeeded, and every one was judged on the system's own tool-call log. No inference, no false positives, no argument about whether it really happened.

Evidence

Findings you can defend, not just read.

Every attack and every response is captured, timestamped and sealed into a hashed manifest. Any later change to any evidence file breaks the hash, and our reporting tool refuses to produce a report from evidence that does not verify. If a finding is ever disputed, by your vendor, your board or an auditor, the record is verifiable rather than merely asserted.

Testing runs on Australian infrastructure. Evidence is stored and destroyed in Australia. Where a model assists our analysis, it is restricted by policy to Australian regions, and we verified that rather than assuming it.

Why now

The AI shipped first. The testing didn't.

Australian organisations put chatbots, copilots and agents into production faster than security could keep up. Attackers are already probing LLMs in the wild, and auditors, boards and enterprise buyers have started asking for AI-specific assurance. The gap between "we deployed it" and "we tested it" is where the risk sits. That gap is what we close.

Start here

Tell us what you've deployed.

We'll tell you how we'd try to break it, and what a first engagement would cover. No obligation, no pressure.

Book a scoping call Or see a sample report →