Why we don't test our own AI
We build AI automation for regulated Australian organisations. Which makes us exactly the wrong people to tell you whether what we built is safe. That isn't false modesty. It's how assurance works.
The short version
The team that builds an AI system knows where its guardrails are, and that knowledge blinds them to the ways around them. You test the paths you built. An attacker doesn't.
AI red teaming is adversarial testing of a deployed system: what it can be talked into, coaxed to disclose, or made to do with its own tools. Your pen test doesn't cover it.
So we don't test our own work. When it needs independent testing, we point clients to a separately-run red team that deliberately doesn't test what we built.
Australian organisations put AI into production faster than anyone tested whether it could be trusted. Chatbots on the front page, copilots inside the business, agents wired to real tools with real permissions. The systems shipped. The testing didn't. We build these systems, so we say this from the inside: building carefully is necessary, and it is not the same thing as proving the result holds.
What AI red teaming actually is
Red teaming is adversarial testing. You take the system you've deployed and attack it the way a motivated outsider would, on purpose, under authorisation, and you document what breaks. Not a checklist. Not a scan. A person deliberately trying to make your AI do something it shouldn't.
It matters because the failure mode of an AI system isn't the failure mode of ordinary software. Traditional security asks whether someone can break in. Red teaming an AI asks a stranger question: what can the system be talked into. Can it be made to ignore its own instructions. Can it be coaxed into revealing data it was never meant to show. Can its connected tools be turned against the business. The system doesn't need to be hacked in the old sense. It just needs to be convinced.
Why your pen test missed it
A penetration test checks the network and the API. It's necessary, it's mature, and it does not look at the model. The attack surface that opened when you added an LLM, prompt injection, guardrail bypass, system prompt leakage, tool and function abuse, excessive agency, sits underneath the layer your pen test covers. The OWASP Top 10 for LLM Applications exists precisely because this is a distinct category of risk. Most organisations testing AI today are testing everything around it and nothing inside it.
Why the team that built it can't be the team that tests it
This is the part that matters most, and it's structural, not a criticism of anyone's competence. The people who built the system know where the guardrails are. That knowledge is exactly what stops them finding the ways around them. They test the paths they designed. An adversary doesn't care about the design, they care about the gap between what you intended and what you shipped, and that gap is invisible from the inside.
Most AI in production was built by a capable team, an in-house group or an external vendor who knew what they were doing. That's not the problem. The problem is that none of them can be the ones to test it. A vendor won't adversarially attack the system they just sold you. An internal team can't find the gaps in a design they hold in their heads. It doesn't matter how good the builder is. The test has to come from outside the build, or it isn't a test.
Which is why we hold the line on our own work. Assurance of your own build isn't assurance, it's marking your own homework, and we won't hand a client a document that pretends otherwise.
The scanner isn't the answer either
There are automated tools that promise to red team AI for you. They have a place, and they also lie confidently. On one engagement, an industry-standard jailbreak probe reported a 99.6% attack success rate across more than a thousand attempts. A board reading that number would have concluded the system was completely broken. The actual responses, read by hand, told the opposite story: not one showed the system breaking role. It had defended itself on every attempt and been scored as a total failure.
That's the work. The tools generate noise. Someone with judgement reads the noise and tells you what actually happened, against the system's own logs, not a scanner's guess.
So who does the testing
When the work we build needs adversarial testing, and increasingly, for regulated buyers, it does, that testing has to come from outside the build. There's a separately-run business we point clients to for it: Provok, an independent AI red team, onshore and authorised, which deliberately does not test systems we've built. Same reason we won't test our own: the moment the tester and the builder are the same hand, the test stops meaning anything.
If you're deploying AI that matters, build it carefully, then have someone who didn't build it try to break it. We can do the first part. We'll tell you honestly who does the second.
Building AI that has to hold up?
We build AI automation for regulated Australian teams, onshore, documented, with human oversight designed in. Request a call and we'll tell you straight what's right for your situation, including when you need independent testing we won't do ourselves.