As artificial intelligence has taken off in earnest, the spotlight is more and more on the data that is being used to test software applications and other digital infrastructure.
The quality of data becomes increasingly a key focus in QA, as clean data serves as the foundation for any successful AI application.
AI algorithms learn from data; they identify patterns, make decisions, and generate predictions based on the information they’re fed.
Consequently, the quality of this training data is paramount, stressed Anthony Deighton, a seasoned veteran in the enterprise software industry and general manager of data products at Tamr, overseeing Tamr’s product and solutions strategy.

“Poor testing data quality can come in various forms,” he explained. “From incomplete data with missing fields and inconsistent data with mismatched formats to irrelevant data that does not align with the business’s objectives.”
Massachusetts-based Deighton said that “when such data is fed into an AI system, the consequences can range from mild inaccuracies to severe operational disasters.”
Incorrect predictions and testing could lead to flawed strategic decisions, while biased algorithms could result in reputational damage and legal issues, he pointed out.
“Therefore, prioritising strategies for creating clean training data is crucial for organizations to harness the full potential of AI technology,” Deighton wrote in a recent analysis.
The role of AI
While the problem of data quality may seem “daunting,” as Deighton put it, he does stress “there is hope.”
In fact, “the very technology affected by data quality, AI, can also play a pivotal role in enhancing it,” he explained, because AI-powered automated data cleaning tools can detect and rectify anomalies in the data.
“These tools can identify missing data, spot inconsistencies, and effortlessly remove redundant entries, providing a single, accurate view of each data point,” Deighton continued.
“At the heart of the AI revolution, data quality becomes the master key that unlocks AI’s full potential.”
– Anthony Deighton
Furthermore, they excel in data unification, seamlessly merging and reconciling data from disparate sources into a cohesive, user-friendly format. AI transforms data cleaning from a daunting task into a streamlined, automated process.
“Human review of the data surfaced by AI’s advanced algorithms is crucial in creating quality training data,” he said.
“Human intelligence effectively guides AI in curating data for optimal output.”
The partnership between AI and human expertise ensures that the training data fed into AI models is of the utmost quality, Deighton stressed, resulting in more robust and accurate AI systems.
“By embracing AI with human feedback in their data management strategy, organizations can maintain high-quality data, substantially boosting their AI systems’ performance.”
Pitfalls
The best way to avoid the pitfalls of poor testing data is to ensure its quality from the outset, Deighton noted. This is where data products come in.
“But there’s often confusion surrounding the term ‘data product,’ leading to various interpretations of the definition,” he said.
“To bring some clarity to the discourse, a data product is a consumption-ready set of high-quality, trustworthy, and accessible data that people across an organization can use to solve business challenges.”
Organised by business entities and governed by domain, data products are the best version of data.
Deighton explained they are comprehensive, clean, curated, continuously-updated data sets, aligned to key entities such as customers, vendors, or patients, that humans and machines can consume broadly and securely across an enterprise.
“Data products, powered by AI-driven efficiency with human oversight to provide feedback, play a crucial role in the collection and management of data, guaranteeing its quality and reliability,” he said.
Deighton continued: “At the heart of the AI revolution, data quality becomes the master key that unlocks AI’s full potential.”
In the pursuit of data quality, AI-powered data products emerge as the solution, ensuring accuracy and reliability.
“Investment in data quality isn’t a discretionary business decision—it’s an essential commitment to the future of AI-enabled innovation,” Deighton concluded.
“The key to avoiding the trap of ‘garbage in, garbage out’ lies not in the sophistication of your AI, but in the quality of your data,” he summarised.
NEXT WEEK IN SINGAPORE

REGISTRATION IS NOW OPEN FOR THE QA FINANCIAL FORUM SINGAPORE 2024
Test automation, data and software risk management in the era of AI
The QA Financial Forum launches in Singapore on November 6th, 2024, at the Tanglin club.
An invited audience of DevOps, testing and quality engineering leaders from financial firms will hear presentations from expert speakers.
Delegate places are free for employees of banks, insurance companies, capital market firms and trading venues.

QA FINANCIAL FORUM LONDON: RECAP
Last month, on September 11, QA Financial held the London conference of the QA Financial Forum, a global series of conference and networking meetings for software risk managers.
The agenda was designed to meet the needs of software testers working for banks and other financial firms working in regulated, complex markets.
Please check our special post-conference flipbook by clicking here.
READ MORE
- Cognizant drags rival Infosys to court over trade secrets
- Testaify claims tool is ‘100x faster than seasoned QA architect’
- Fast-growing Newgen sets sights on banks in Middle East
- ABN Amro hires nCino and CBA for digital upgrade
- QAFF London: Lloyds’ Richard Bishop on the rise of ‘green software’
Become a QA Financial subscriber – for FREE
* Receive our weekly newsletter * Priority invitations to our Forum events


