Deep Dive: why AI’s limits matter for software testing in financial services

Atlanta-based Ashlee Gardner
Atlanta-based Ashlee Gardner

Artificial intelligence is no longer a novelty, it’s at the center of how work is being transformed across industries, especially in highly regulated sectors like finance.

But as buzz grows around artificial general intelligence (AGI), software testers and QA teams must cut through the hype and understand what today’s AI can, and cannot, do. A recent deep dive from Emory University offers exactly that clarity.

Speaking on behalf of Emory’s Center for AI Learning, Atlanta, Georgia-based Ashlee Gardner, Assistant Director of Strategic Communications, explained that “artificial intelligence is everywhere lately, on the news, in podcasts and around every water cooler. A new, buzzy term, artificial general intelligence (AGI), is dominating conversations and raising more questions than it answers.”

As the conversation expands from general AI to AGI and even artificial super intelligence (ASI), it’s crucial for engineering teams, especially those building or deploying automated testing frameworks, to understand what’s real today and what remains speculative.

“The definitions have changed over time, and that has caused some confusion. Artificial intelligence is not a monolithic technology. It’s a set of technologies that automate tasks or mimic decisions humans normally make,” Gardner explained.

This is particularly relevant to QA professionals in financial services, where AI-driven test automation tools are becoming commonplace.

While these tools offer efficiency, Gardner highlighted the critical distinction between narrow AI and AGI: “Traditionally, when people talked about artificial general intelligence, they meant Skynet from ‘The Terminator’ or HAL from ‘A Space Odyssey’, machines that supposedly had free will and approximated human abilities.”


“We need thorough testing to identify where these models break or produce biased outcomes.”

– Ashlee Gardner

Today, major research labs have started redefining AGI to mean AI that performs as well as, or better than, expert humans at specific tasks. But Gardner cautioned against this loose use of the term.

“Large language models like ChatGPT can outperform humans trying to get into medical school on the MCAT. But that’s not real intelligence,” Gardner stressed.

“It is like giving a student Google during an exam. True AGI should show reasoning, not just information retrieval and pattern matching.”

In QA workflows, that distinction matters. Automated systems that seem ‘smart’ may simply be reproducing learned outputs, not understanding systems behavior, risk exposure, or compliance nuances.

“Today’s models give the impression they are reasoning, but they’re just sequentially researching information and then summarizing it,” she explained. “They don’t understand the world, they just predict what word comes next based on patterns.”

This misunderstanding of AI capabilities has implications for how testing teams evaluate AI-powered tools. Gardner offered a simple but striking example: “When tested on real reasoning tasks, like the Tower of Hanoi or logic puzzles, LLMs often fail unless they’ve memorisMed the answers.”

Even humor, something as seemingly trivial as it is profoundly human, exposes AI’s limitations. “From a humanities perspective, humor lies at the intersection of comfort and discomfort. That boundary shifts all the time. Chatbots only regurgitate things they’ve seen in the past; they don’t understand that boundary,” she shared.

The same pattern applies to decision-making in business contexts. Most enterprise AI tools are trained on internal datasets, making them effective at historical analysis but often incapable of synthesizing forward-looking insights.

“Their AI models, which are trained on internal data, can’t synthesize where we’ve been with where we’re going. That would require reasoning, intuition and values alignment, things we struggle to articulate even for ourselves,” Gardner said.

Human oversight

From a software testing perspective, this underscores the need for robust human oversight.

While automated systems can efficiently flag failures or generate tests, they are not yet capable of adaptive reasoning or values-based judgment.

Gardner was clear on this point: “LLMs won’t bring us anywhere close to AGI because they don’t have reasoning or intuitively creative abilities. We haven’t given them a framework to efficiently discover new information.”

That framework may be on the horizon in research settings. “One step in the right direction is joint embedding predictive architecture, or JEPA,” she said. “Instead of stringing words together like an LLM, it infers deeper relationships between concepts and activates those inferences to achieve a higher-level objective.”

In the meantime, Gardner offered reassurance that there’s more to human intelligence than information compression. “It’s refreshing to learn that throwing all of society’s encyclopedia entries into an LLM doesn’t produce human-level intelligence. I’m boiling it down, of course. There’s more to humanity than meets the eye.”

Still, current AI technologies are undeniably powerful when used appropriately.

“They enable people to do tasks in hours that they previously had to spend days on. The promise is efficiency, tools that can summarize research, assist in medical diagnosis, help you plan your shopping list for the week,” Gardner said.

This promise aligns directly with the goals of QA automation: speeding up cycles, improving accuracy, and freeing up human engineers for more complex, meaningful work.

But Gardner warned that without public understanding, that promise could quickly become peril. “The peril isn’t the tech; it’s the lack of public understanding. We need AI literacy, so people understand when AI is being used the right way and when it is not.”

Testing teams, especially those working in regulated financial environments, should take this as a cue to approach AI tooling with both optimism and skepticism. “We need thorough testing to identify where these models break or produce biased outcomes,” she emphasised.


“LLMs won’t bring us anywhere close to AGI because they don’t have reasoning or intuitively creative abilities.”

– Ashlee Gardner

Oversight and transparency are not just technical requirements, they are social obligations. Gardner said. “Some companies have argued in court that scraping people’s data without consent or compensation is justified because it advances society. That’s a manipulative and troubling argument.”

She suggested a more constructive approach: “If the needs of the people contributing to these technologies are represented and they are adequately rewarded, it would incentivise greater innovation and usage.”

On the more philosophical side, Gardner emphasised that aligning AI with human values isn’t just a technical challenge, it’s a social one.

“One audience member asked, ‘How can we ensure that the AI models we are building have American values? What makes America special is that we value the opportunity for dissent,” she said. “We’re never going to agree on everything, so the bigger question is: How do we create a system with responsive guardrails that adapt to our evolving societal values?”

These are not new questions, but the speed of AI development is forcing a fresh conversation around collective responsibility and transparency in AI design. And for the software testers and engineers helping deploy these systems, the implications are direct.

“If we get AGI right, what does the best version of the future look like?” Gardner asked.

Her answer is simple but powerful: “I think the best version of the future with all these technologies is one where they free us to do more meaningful work and spend less time on things we don’t enjoy.”

She pointed to Emory’s ongoing AI research as an example of that future in action. “Our research is already transforming health care, leading to improved diagnosis and treatment of diseases like cancer, heart disease and diabetes. We’re using text analysis to uncover patterns in public policies that will make governance more efficient and equitable. Our scholars are also looking at how AI can protect people’s rights and grow businesses.”

Ultimately, Gardner’s message is one of human-centered optimism.

“AI offers more benefits than drawbacks, if we empower people through education and include them in the conversation so they can advocate for themselves and those they care about. If deployed thoughtfully, these technologies can amplify human potential, not replace it.”


NEW EVENT


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

REGISTER HERE TODAY



REGULATION & COMPLIANCE

Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.


READ MORE


WATCH NOW