Provok Book a scoping call
How we work

Rules before results.

Letting someone attack your production AI is a risk in itself. These are the rules we work under, what we test against, how long it takes, and what happens to your data. If any of it does not suit your organisation, tell us at scoping and we will adjust or decline the work.

Rules of engagement

What we will never do.

These are not preferences. They are enforced in our tooling, and an engagement cannot run without them.

Never

Test without written authorisation

Every engagement begins with a signed authorisation record naming the system, the endpoint, the environment and the test window. Our testing platform refuses to run until that record exists and is confirmed. No authorisation, no traffic. That is a technical control, not a promise.

Never

Touch production without you asking us to

We test staging or a dedicated test environment by default. If you need production tested, that has to be explicitly authorised in writing, scoped to a window you choose, and we will tell you plainly what the risks are first.

Never

Use real customer data

We work with synthetic accounts and test records. If your environment contains real personal information, we say so at scoping and either have it removed from scope or decline the engagement. Finding a vulnerability is not worth creating a breach.

Never

Sit on something critical

If we find something serious mid-engagement, you hear about it that day, not in the report three weeks later. We stop, we tell you, and we agree together whether to keep testing.

Never

Send a finding we have not verified ourselves

Automated AI security tools produce false positives at scale. Every flagged result is checked against the system's actual response or its actual tool-call log before it becomes a finding. Nothing reaches your report on a scanner's word alone.

Never

Take your data offshore

Testing runs on Australian infrastructure. Evidence is stored and destroyed in Australia, under the retention period recorded in your engagement agreement. Where we use a model to assist analysis, it is restricted by policy to Australian regions.

Coverage

What we test against.

We map findings to the OWASP Top 10 for LLM Applications, so your report speaks the language your auditors, your board and your enterprise customers already use. Not every category applies to every system. Scoping decides which are in play.

RefCategoryWhat we do about it
LLM01Prompt InjectionDirect and indirect injection, including instructions hidden in documents, tickets and other content the system ingests.
LLM02Sensitive Information DisclosureAttempts to extract system instructions, credentials, internal notes and other customers' records.
LLM03Supply ChainReviewed at scoping. We flag model, plugin and integration exposure, but we do not test third-party vendors without their authorisation.
LLM04Data and Model PoisoningWhere you control a knowledge base or retrieval corpus, we test whether content placed into it changes the system's behaviour.
LLM05Improper Output HandlingWhether model output is passed downstream without validation, into rendering, queries, or other systems.
LLM06Excessive AgencyWhether the system can be made to take actions beyond its remit, or on behalf of a user it has not verified.
LLM07System Prompt LeakageExtraction of the instructions, rules and context the system was configured with.
LLM08Vector and Embedding WeaknessesCross-tenant retrieval, and whether one user's documents can surface in another user's answers.
LLM09MisinformationWhether the system fabricates facts, accepts false premises, or asserts capability and compliance it cannot support.
LLM10Unbounded ConsumptionWhether an attacker can drive cost, exhaust quota, or degrade the service for legitimate users.

We also test things the framework does not name, including multi-turn escalation, session and context isolation between users, and business logic specific to what your system is for. A framework is a floor, not a ceiling.

Process

What actually happens, and when.

Indicative timings for a single-system engagement. Larger scopes take longer, and we tell you the real number at scoping rather than after you have signed.

Day 0

Scoping call

Thirty minutes. What you have deployed, what it connects to, what would hurt if it failed. You leave knowing whether testing is worth it for your setup, and roughly what it would cost. No obligation.

Day 1 to 3

Paperwork and authorisation

Engagement agreement, your NDA or ours, and the signed authorisation record naming the exact system and window. Nothing runs until this is complete.

Test window

Attack

Broad automated coverage first to map the surface, then targeted and adaptive attacks where the surface looks weak. Anything critical is reported to you the same day it is found.

Then

Verification

Every flagged result is read by a human against the system's actual behaviour. Tool noise is discarded. Only what genuinely happened becomes a finding, with a severity and a framework mapping we set, not a scanner.

Within a week

Report

Findings with evidence, business impact, and remediation your engineers can act on. Plus what we could not test and why, because coverage gaps matter as much as findings. A walkthrough call is included.

After you fix

Retest

We re-run the same attacks and confirm each finding is closed, or not, with fresh evidence. Included in the engagement, not an upsell.

Your data

Confidentiality and evidence handling.

Confidentiality. We will sign your NDA, or provide ours, before scoping goes into any detail. We do not name clients publicly without written permission. Sample reports on this site are either anonymised or drawn from our own test systems.

Evidence. Every attack and every response is captured, timestamped and sealed into a hashed manifest. Any later change to any evidence file breaks the hash, and our reporting tool refuses to produce a report from evidence that does not verify. Your report includes the manifest so your own team can confirm nothing was altered after sealing.

Location and retention. Evidence is stored in Australia for the retention period recorded in your engagement agreement, then destroyed. You can request earlier destruction at any point after the report is delivered.

Access. Engagement evidence is held on a dedicated, network-restricted machine used only for testing. It is not stored on general-purpose devices and is not shared with third parties.

Full detail is on our security and data page, and the specifics for your engagement are written into your agreement rather than left to policy.

Want to know how your AI holds up?

Tell us what you've deployed. We'll tell you how we'd try to break it, and what a first engagement would cover. No obligation.