Provok Book a scoping call
Systems we test

If it takes input,
it can be turned.

Different AI systems fail in different ways. A chatbot that says the wrong thing is embarrassing. An agent that takes the wrong action moves money. Here is what we test, what actually goes wrong, and where each one starts.

Customer-facing chatbots

From $3,500
What it is

The support assistant, sales bot or FAQ widget on your website or in your app. Anything a member of the public can type into.

What goes wrong

It leaks its own instructions and any secrets embedded in them. It gets talked past its guardrails into advice, commitments or language it was never meant to give. It reveals what it knows about your business, your pricing, or other customers.

Why it matters

It is the most exposed AI you own. Anyone on the internet can attack it, and anything it says, your organisation said.

RAG and knowledge-base assistants

From $6,500
What it is

An assistant that answers from your documents: policies, claims files, member records, contracts, clinical guidelines. If you connected an AI to your own knowledge, this is you.

What goes wrong

Data it should never surface. We test whether a user can be shown another customer's record, an internal-only document, or content outside their permission level, and whether the retrieval layer can be manipulated into fetching the wrong thing.

Why it matters

This is a Privacy Act problem, not an embarrassment. In health and financial services, a cross-customer disclosure is a notifiable event.

Agentic AI with tools

From $9,500
What it is

AI that does things, not just says things. It looks up records, triggers workflows, sends messages, updates systems, processes requests. Anything wired to your business through tools, functions or automation.

What goes wrong

The agent is manipulated into taking an action it should not, for someone it should not act for. Tools invoked without verifying identity. Actions triggered by instructions hidden in content the agent read. Permissions far broader than the conversation warranted.

Why it matters

The highest-consequence AI in most organisations, and usually the least tested. A finding here is not "it said something odd", it is "an attacker made it act".

Internal copilots and staff assistants

From $5,500
What it is

The assistant your own people use. Drafting, summarising, answering internal questions, helping developers write code.

What goes wrong

Staff-facing systems are usually the least governed, because "it's only internal". We test for data crossing team boundaries, secrets surfacing in responses, and coding assistants producing insecure output or exposing repository contents.

Why it matters

Internal does not mean safe. It means unmonitored, and a compromised internal assistant is a foothold.

Voice agents

From $6,500
What it is

Phone-based and spoken assistants handling enquiries, bookings, triage or claims.

What goes wrong

The same failures as any assistant, with real-world consequences: identity assumed rather than verified, actions taken on a caller's word, information disclosed to whoever is on the line.

How we test it

We test the model and tool layer beneath the voice, where the decisions are actually made and where the vulnerabilities live. Testing through the speech pipeline itself is scoped separately if you need it. We will tell you plainly which you are getting.

Document and email AI

From $6,500
What it is

AI that reads content it did not write: inbound email, uploaded documents, tickets, forms, web pages.

What goes wrong

Indirect injection. Instructions hidden inside a document or email that the AI obeys as though you had typed them. The attacker never speaks to your AI directly, they just send it something to read.

Why it matters

The most underestimated attack surface in AI today. Every uploaded file and inbound message is untrusted input, and almost nobody treats it that way.

Not sure which one you have?

Most organisations have more than one, and often did not realise the second existed. Tell us what you have deployed on a short scoping call and we will tell you what is worth testing, what it would cover, and a fixed price before anything begins.

Every engagement includes a severity-rated findings report, remediation direction your engineers can act on, framework mapping, and a retest once you have fixed what we found. Prices start where shown and are confirmed on scope. See how engagements are structured or what the report looks like.

Book a scoping call →