Software testing teams at banks and financial institutions may need to rethink many of the principles that have underpinned quality assurance for decades as agentic AI moves from experimentation into production.

Unlike conventional enterprise software, autonomous AI agents can make different decisions when presented with similar circumstances, introducing a level of unpredictability that traditional testing frameworks were never designed to evaluate.

As banks increasingly deploy AI across software development, operations and customer-facing processes, the challenge is becoming less about verifying whether software functions correctly and more about proving that autonomous systems behave safely, consistently and within defined governance boundaries.

For quality engineering teams, that represents a significant shift. Functional testing, integration testing and performance validation remain essential, but they no longer provide sufficient assurance for systems that continuously adapt their behaviour according to context.

Instead, software testing increasingly needs to evaluate the quality of AI decision-making itself, including explainability, confidence levels, policy compliance and resilience when conditions change.

“Traditional QA validates execution. Agentic AI quality engineering must validate behavior.”

– Gaurav Aggarwal

That evolution is particularly relevant for financial services, where regulations such as DORA and wider AI governance initiatives place growing emphasis on operational resilience, auditability and evidence.

For autonomous software, demonstrating how a decision was reached may become just as important as demonstrating that the final outcome was correct.

Gaurav Aggarwal, global head of solutions engineering at California-based technology consultancy WinWire, argued that this represents a fundamental break from previous generations of software testing.

Gaurav Aggarwal

“Agentic AI breaks that assumption,” he stated, referring to the long-held belief that software behaves predictably when given the same inputs.

“It doesn’t just challenge the tools we use, it invalidates the mental model quality teams have historically relied on.”

Aggarwal said many organisations are already discovering that autonomous systems can successfully pass traditional testing while still failing to generate sufficient trust for production deployment.

“The agents often work as intended, yet releases slow down. Not because tests fail, but because teams struggle to answer a more fundamental question: Can we trust the decisions these systems are making?”

He stressed “that is where traditional quality engineering reaches its limits” as Aggarwal pointed out that conventional quality engineering focuses on validating correctness, predefined workflows and expected outcomes, while autonomous AI introduces probabilistic behaviour that requires an entirely different testing mindset.

“Behavior is probabilistic. Decisions evolve with context. Failures may not repeat in the same way. Quality must therefore extend beyond outcomes to include intent, confidence and boundary adherence, ” he wrote in Forbes Technology Council.

That, Aggarwal continued, changes the very purpose of software testing. “Put simply, traditional QA validates execution. Agentic AI quality engineering must validate behaviour.”

For banks deploying AI into business-critical systems, he argued that software quality can no longer be reduced to pass or fail.

“Quality engineering must evolve faster than code; otherwise, agentic AI will move quickly, learn rapidly and fail expensively.”

– Gaurav Aggarwal

Instead, testing teams need to determine whether an AI agent made the right decision for the situation, whether it remained within approved operational limits, whether its reasoning can be explained, and whether it can safely recover when circumstances change.

“In regulated and high-stakes environments, these are audit and liability requirements,” Aggarwal shared. “A system that produces a correct outcome for the wrong reason still represents a quality risk. In agentic systems, how a decision is made matters as much as the result.”

Aggarwal also warned against viewing AI testing simply as using AI tools to test other AI systems.

“The main problem is figuring out what quality means for systems that learn and change. Tools might help, but trust is built by being clear about your intentions, setting clear boundaries and having someone in charge who is responsible.”

He argued that quality engineering is rapidly becoming a strategic concern rather than simply an engineering discipline. “Quality engineering for agentic AI is not just a technical concern, it is a leadership responsibility.”

Ultimately, Aggarwal believes organisations that succeed with agentic AI will be those able to continuously demonstrate that autonomous systems behave responsibly throughout their lifecycle, rather than simply deploying AI faster than competitors.

“Quality engineering must evolve faster than code; otherwise, agentic AI will move quickly, learn rapidly and fail expensively.”

He concluded: “The future belongs to organisations that treat quality not as a checkpoint but as a discipline purpose-built for intelligent, autonomous systems.”

REGISTER TODAY – SIMPLY CLICK HERE

Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY

REGULATION & COMPLIANCE

Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.

READ MORE

WATCH NOW

QA FINANCIAL PODCASTS