It has been estimated that more than 80 per cent of AI projects fail, which is roughly twice the failure rate of comparable IT projects that do not involve AI.
Additionally, research by MIT has found that 95 per cent of organisations are seeing no measurable return from their generative AI pilots. Only a small minority, around five per cent, are extracting real value at scale.
While there are often a number of factors behind these failures, from our experience, the standout reason is frequently a lack of quality data for AI to work with, rather than the technology itself.
Data quality problems and ongoing decay
Currently, 94 per cent of organisations estimate that they have data quality issues, and are operating with inaccurate, duplicate or incomplete data. This is driven by data decay which is a significant factor impacting on the effective implementation of AI.
Customer contact data can decay at around 25 per cent annually as people move home, die and get divorced. Additionally, 20 per cent of addresses entered online contain errors, including spelling mistakes, incorrect house numbers and wrong postcodes. This often results in data teams spending the majority of their time on data preparation – cleaning and structuring data so it’s ready for analysis – rather than generating insights. For those that don’t make any effort to clean their data they should expect their AI efforts to fail.
After all, access to clean data is essential for organisations looking to train, deploy and scale AI, as well as determine the return on investment (ROI) from their AI initiatives. Inaccurate data can result in unreliable automation, ineffective personalisation, poor recommendations, and ultimately, a loss of customer trust.
Improve data verification processes
To tackle the problem of inaccurate contact data, I strongly recommend putting verification processes in place both at the point of data capture and when cleaning existing data in batch. In many cases, this can be achieved through simple, cost-effective improvements to the data quality process.
Use address lookup or autocomplete
To begin with, I recommend using an address autocomplete or lookup service at the customer onboarding stage. These services deliver accurate address data in real-time by supplying a properly formatted and correct address as customers begin entering theirs. They can also reduce the number of keystrokes needed to type an address by up to 81 per cent. This results in a speeded up onboarding process, and also reduces the likelihood of a customer abandoning an application or purchase. It’s also important to extend this approach to the first point of contact verification for email and phone numbers, ensuring these valuable contact data channels are verified in real-time.
Prevent data duplication
A significant issue is data duplication because customer databases commonly have duplicate rates of 10 –30 per cent. This commonly occurs when two departments merge their data or when errors in contact data collection take place at different touchpoints. Duplication can confuse an AI application, while also adding cost in terms of time and money, particularly when printed communications are involved, which can negatively impact on the sender’s reputation. I recommend using an advanced fuzzy matching tool to deduplicate data. This type of service can merge and purge even the most challenging records to create a ‘single user record’, delivering an optimum single customer view (SCV) from which AI can make learnings.
Undertake data cleansing / suppression
Data suppression, or cleansing, are other important parts of the data cleaning process. Using such services it’s possible to identify people who have moved or are no longer at the address on file. This is a vital element of data cleaning and, consequently, of supporting AI initiatives. As well as removing incorrect addresses, these services can include deceased flagging to prevent mail and other communications being sent to people who have passed away, which can cause distress to their friends and relatives. Employing suppression strategies can help organisations to save money, protect their reputations, avoid fraud and improve the quality of the data supporting their AI efforts.
Enrich data for AI success
It’s equally important to enrich data with demographic, firmographic, geographic, social media and property attributes, as well as add missing email and phone information to support and maximise AI’s efforts in analytics, personalisation and omnichannel marketing. A more complete data picture will deliver increasingly accurate predictions by giving AI additional signals to identify patterns, predict customer needs and assess likely outcomes, resulting in improved recommendations and communications.
Ensure data is machine readable
AI agents need to be able to draw from a foundation of high quality, API or machine readable data. Because AI systems need data they can access, interpret and process quickly at scale without relying on manual intervention.
It’s structured, well formatted machine readable data that reduces ambiguity and makes it easier for AI models to interpret information correctly. At the same time decision making is speeded up as AI can analyse and act on machine readable data much faster than humans can. Also, it improves integration asmachine readable data can flow more easily between different systems, applications and AI tools.
Therefore, it’s strongly recommended to use well labelled data to ensure accuracy across mission critical AI applications, because high quality data labelling helps to maintain accuracy, consistency and trust as AI systems grow.
In summary
Having robust processes in place to deliver high quality data to maximise success from AI efforts is essential, because it can deliver significant benefits, in particular in driving profitability. Failing to do so can result in AI generated hallucinations and poor outcomes, while slowing the adoption of valuable AI technology within an organisation.
Barley Laing
Barley Laing established and leads the UK office of US-headquartered global data quality and ID verification business, Melissa.
As Managing Director, with 29 years of technology and data industry experience, his role is focused on meeting the data quality, address and ID/compliance needs of organisations in the UK, Ireland, Scandinavia, and beyond.


