Everyone is talking about AI. Organisations are investing in copilots, automation platforms, predictive analytics, and generative AI at an incredible pace. But there's one conversation I still don't think we're having enough.
Your AI is only ever going to be as good as the data you give it.
I spend my career helping organisations clean messy data, and one misconception I encounter time and time again is the belief that AI can somehow fix poor-quality data automatically. It can't.
In many cases, it actually makes the problem worse.
If your customer appears as "Microsoft Ltd", "Microsoft Limited", "MSFT" and "Microsoft Inc." across the same dataset, an AI model may infer that they're the same company—or it may not. If dates are stored in multiple formats, records are duplicated, or important fields are missing, AI isn't removing those inconsistencies. It's making decisions based on incomplete or conflicting information.
The result is unreliable reporting, inaccurate analytics, and automation that people simply can't trust.
That's exactly why I created my new five-day LinkedIn Learning challenge. Before organisations think about AI readiness, they need to think about data readiness.
Stop Asking AI to Guess
One of the biggest mistakes I see is people jumping straight into AI tools without understanding the quality of the information they're feeding them.
We expect AI to identify duplicates, interpret abbreviations, understand inconsistent naming conventions, and work around poor formatting.
Sometimes it can.
But should it have to?
Every time AI has to make an educated guess, you're introducing another opportunity for error. Those errors quickly compound when you're dealing with thousands—or even millions—of records.
Good data removes the guesswork.
My COAT Framework
Whenever I assess a dataset, I use what I call the COAT framework. It gives me a simple way to determine whether data is genuinely ready for analytics, automation, and AI.
COAT stands for:
Consistent – Is the same information recorded the same way every time?
Organized – Is the data structured logically and easy to understand?
Accurate – Does it reflect reality?
Trustworthy – Can you rely on it to support business decisions?
These four principles sound simple, but they're surprisingly effective at uncovering the issues that quietly undermine analytics projects.
Before fixing anything, I always encourage people to audit their data against these four questions. You can't improve what you haven't identified.
When Data Strategy Leads AI
Why leaders must treat data quality, ownership and governance as the first phase of AI delivery, not a back-office cleanup task.
Cleaning Before Automating
Once you've identified the problems, it's time to standardise your data.
Many of the issues I encounter don't require complex software. Simple Excel functions like TRIM, UPPER, LOWER, PROPER, and Find and Replace can eliminate a huge number of inconsistencies in just a few minutes.
This is where AI becomes genuinely useful.
Rather than asking AI to clean your data for you, ask it to help you clean it.
It can generate formulas, recommend naming conventions, build lookup tables, identify similar records, or suggest standardisation rules. Used this way, AI becomes an incredibly valuable assistant.
But there's one thing AI still doesn't have.
Context.
It doesn't understand your customers, your suppliers, your internal terminology, or your business rules as well as you do. Human judgement is still essential.
Structure Matters Just As Much As Accuracy
Even after you've cleaned your data, there's another challenge that often gets overlooked.
Structure.
I regularly see spreadsheets where multiple pieces of information are stored in a single cell. Contact details, company names, countries, job titles—they're all bundled together.
Humans can usually work around that.
AI struggles.
One of the simplest principles I teach is this:
One row should represent one record. One column should represent one attribute.
When every column contains just one type of information, analytics become easier, reporting becomes cleaner, and AI has far less interpretation to do.
You're reducing ambiguity before your AI tools even begin processing the data.
Data Governance as Architecture
Treat information policies as core design, creating resilient structures that support insight, compliance, and long-term agility.
Data Readiness Is AI Readiness
The purpose of my LinkedIn Learning challenge isn't simply to teach people how to clean spreadsheets.
It's to change how they think about preparing data for AI.
Over five days, learners work through a realistic dataset, auditing its quality, standardising inconsistencies, restructuring information into an AI-ready format, validating the results, and ultimately producing a dataset they can genuinely trust.
By the end of the challenge, they don't just have cleaner data—they have a repeatable framework they can apply to almost any dataset they encounter.
The Future of AI Starts With Better Data
As AI becomes embedded in more business processes, the quality of the underlying data will become an even greater competitive advantage.
The organisations seeing the best AI outcomes won't necessarily be those with the biggest models or the newest tools.
They'll be the ones with the cleanest, most trustworthy data.
Because AI isn't a magic wand.
It's a multiplier.
If your data is organised, accurate, and reliable, AI can help you achieve remarkable results.
If it isn't, AI will simply help you make bad decisions faster.
That's why I believe data quality isn't just part of AI readiness.
It is AI readiness.
Comments ( 0 )