IT often begins with a brilliant idea – one that promises a step change in productivity and, in turn, competitive edge. A team builds a prototype using artificial intelligence (AI), early demonstrations look promising, and business leaders are keen to move fast. But then, momentum stalls.
Governance teams raise red flags and product managers question whether the model is robust enough for external use. Meanwhile, engineers struggle to fully explain testing results to the rest of the team. The project doesn’t collapse – but it doesn’t launch either.
These delays reveal a simple truth: building AI is no longer the hard part. Validating it is. As companies race to deploy AI, the real bottleneck isn’t innovation – it’s proof of value. Without rigorous and efficient testing, systems can’t earn trust or scale, ending up in an endless cycle of pilot phases.
When teams aren’t aligned, models don’t move
What this means is that convincing people that the tool will be worth the setup costs is unavoidable. But without a common testing approach and context-specific metrics to assess the strengths and weaknesses of the AI model, many teams default to caution. According to Gartner, at least 30 per cent of generative AI projects will be abandoned after proof of concept by the end of 2025, due to poor data quality, inadequate risk controls, escalating costs, or unclear business value.
BT in your inbox
Start and end each day with the latest news stories and analyses delivered straight to your inbox.
This is where independent testing can play a vital role, introducing an objective criterion that will help teams move past debate and align around facts. It also brings technical, business, and governance perspectives into one conversation, reducing the risk of misunderstandings or delays.
Essentially, it is about providing a neutral voice that helps companies figure out what to test, how much to test, and how to communicate the results.
Building is easy; proving is critical
Thanks to open-source models and off-the-shelf tools, it has never been easier to build functional AI prototypes. But building something that works isn’t the same as proving it’s ready. Without rigorous testing, AI models can behave unpredictably, cause errors, or even endanger lives. For companies, this is a costly exercise that could hurt bottom lines and reputations.
There is no lack of examples. McDonald’s pulled its AI-powered drive-thru system after it repeatedly got orders wrong – adding items customers didn’t ask for and mishearing accents. Air Canada was taken to court after its AI chatbot gave a customer incorrect bereavement refund advice, and the airline was held responsible. Meanwhile, Amazon’s self-driving Zoox vehicles were recalled after a crash in Las Vegas exposed flaws in how the AI handled other drivers.
A 2025 global study by KPMG and the University of Melbourne also found that two-thirds (66 per cent) of employees surveyed said they use AI tools without evaluating the accuracy of responses. And over half (54 per cent) reported making errors because they trusted AI tools without checking the results. These failures don’t just happen at the user level – they begin in development, where assumptions go unchallenged and edge cases go untested.
Independent testing functions as both a mirror and a referee – uncovering blind spots, stress-testing assumptions, and helping teams meet regulatory and governance standards. As AI gets used for all kinds of functions, from verifying personal data at government agencies to helping doctors diagnose medical conditions, ensuring that the technology is robust and properly verified is becoming crucially important.
In other words, testing provides the kind of evidence regulators, customers, and executives increasingly expect: not just that AI works, but that it works safely, reliably, and securely.
Early testing beats late-stage rewrites
Regulations are now already changing to ensure that testing is part of the process of AI deployment.
The European Union’s AI Act will soon mandate rigorous pre-deployment testing for high-risk systems. Singapore’s Model AI Governance Framework, while voluntary, is fast becoming industry standard. Around the world, businesses are being pushed by their customers to prove their models are robust, fair, and explainable – before they’re deployed.
But testing shouldn’t be seen as a final checkbox – something that happens after the model has been built. That’s when problems are most costly to fix. Instead, testing should be designed upfront, as part of the development rather than the deployment process.
Investing in early, structured testing isn’t just a compliance move. It helps avoid expensive rewrites, builds internal confidence, and shortens time-to-market. The earlier teams start testing with clear, standardised criteria, the more they can build trust into the system – before trust becomes a barrier.
The same KPMG global report found that four in five people said they would be more willing to trust an AI system when assurance mechanisms are in place, such as monitoring system reliability, human oversight and accountability, responsible AI policies and training, adhering to international AI standards, and independent third-party AI assurance systems.
Close to three in four (74 per cent) also agreed that they would be more willing to trust an AI system if it is assured by an independent third party.
The future of AI won’t belong to the boldest ideas, but to the most rigorously proven ones.
The writer is co-chief executive officer of Resaro, an independent, third-party AI assurance provider that tests and evaluates mission-critical AI systems globally. Resaro is part of the Global AI Assurance Pilot launched by the AI Verify Foundation and the Infocomm Media Development Authority in February this year.






Leave a Reply