Dirty data and AI: the cheapest mistake a small business can fix first
Most small businesses blame the AI model when the answers come back wrong. The real culprit is usually dirty data, and it is the cheapest readiness gap to fix first.
By Carl Chessum
Dirty data is the everyday business data you already hold (your customer list, your sales records, your operational spreadsheets) when it is full of duplicates, gaps, out-of-date entries and fields that contradict each other. It is the cheapest AI mistake a small business can fix, and the one most likely to sink an AI project before it starts. A person reading a messy record spots the contradiction and asks a question. AI does not ask. It takes whatever it is given and produces a confident answer at scale, which is how a small data problem becomes a large business one.
This is not a fringe risk. 95% of AI pilot programmes fail to create measurable value (MIT NANDA Research, 2025), and in the projects I have audited, dirty data is the first domino more often than any model limitation. The good news is that it is also the cheapest gap to close. You do not need a data warehouse or a data team. You need a spreadsheet, an afternoon, and the discipline to do the work before you buy the tool, not after.
What counts as dirty data?

Dirty data is not one problem. It is five, and most businesses have all of them sitting quietly in systems nobody has opened in a while.
- Duplicates. The same customer entered three times, none flagged, each with a different spelling and a different last-order date.
- Missing fields. The records exist, but the column the AI needs (industry, region, renewal date) is blank on a third of them.
- Conflicting sources. The CRM says one thing, the billing system says another, and there is no written rule for which one wins.
- Stale records. Customers who moved, companies that rebranded, prices that changed two years ago and never got updated.
- Free text doing a structured job. Years of notes typed into a comments box that contradict the tidy fields beside them.
None of these stop a human. We patch the gaps without noticing. They stop AI cold, because the model treats every record as true. What good looks like is laid out on the Data pillar page; the short version is that accuracy, consistency and governance decide everything downstream.
Why does dirty data sink AI projects faster than anything else?
Because AI removes the human who used to catch the problem, and adds speed.
At human scale, a bad record is one bad decision that someone usually notices. At AI scale, the same bad record becomes a thousand bad decisions before anyone checks. That is a large part of why 74% of companies struggle to achieve and scale AI value (BCG, 2024): the pilot works on a clean sample, and then the rollout meets the real data.
Dirty data also rarely travels alone. It arrives with a broken process, because the process is where the bad data gets created. A sales team that can close a deal without completing the fields will leave them empty, and no amount of cleaning fixes a tap that keeps running. That combination is the most expensive failure pattern I see, which is why the Process pillar sits directly next to Data.
How clean does my data need to be before I try AI?
Cleaner than most teams assume, but not perfect. The bar is not spotless data. It is honest data: you know what is good, what is questionable, and who fixes the next problem.
A useful way to place yourself is the four readiness bands we use to score the Data pillar in the audit. Find the row that sounds like your business today.
| Band | What it looks like |
|---|---|
| Critical | Data is siloed, inconsistent or poorly labelled. No governance framework in place. |
| Developing | Some data is centralised but quality varies. Governance is informal and uneven. |
| Progressing | Some data is structured and accessible. Governance exists but is not consistently applied. |
| Ready | Data is clean, well-governed, documented and readily accessible for AI use cases. |
Most established small businesses land in Developing or Progressing, not Critical. That is the honest, workable middle: real gaps, fixable in order. The aim is not to reach Ready across the whole business before you touch AI. It is to get the one dataset your first use case depends on to Ready, and to know where the others stand.
How do I clean my data without a data team?
You do it in a spreadsheet, this week, on the one dataset that matters. Here is the sequence I give owner-operators who have a CRM and not much else.
- Pick the single dataset behind your first AI use case. Not all your data. The customer list, or the product table, or the orders. One.
- Export it to a spreadsheet. A flat copy you can sort, filter and mark up without touching the live system.
- De-duplicate. Sort by email, then by phone, then by name. Merge or flag the repeats, and decide which record is the keeper before you delete anything.
- Fill or flag the critical fields. Find the columns your use case actually needs and either complete them or mark the blanks honestly. A flagged gap is safer than a guessed value.
- Resolve the conflicts. Where two systems disagree, decide which one wins and write that rule down. That decision, not the cleaning, is the part that lasts.
- Standardise the obvious formats. Dates, phone numbers, country names, categories. Pick one format per column and make the column match it.
- Name an owner and a recheck date. Data decays the day after you clean it. Someone has to own keeping it clean, with a date in the calendar.
That is an afternoon for a focused dataset, a few days for a messy one. If the work turns out to be bigger than a spreadsheet (several systems that genuinely cannot agree, or no clear owner anywhere), that is where a data strategy engagement earns its place. Most small businesses do not need that yet. They need the seven steps above, done once, properly.
Start with the gap, not the tool
The cheapest time to find dirty data is before it is feeding an AI tool, not after. The AI Readiness Audit scores your data, and the five other foundations, in about seven minutes. 30 questions, an honest band for each pillar, no card and no sales call.
To see how the data score fits the wider picture, the AI Readiness Audit overview walks through all six pillars and the three ways to run it.
Frequently asked questions
What is dirty data in the context of AI?
Dirty data is the everyday business data you already hold (customer records, sales data, operational spreadsheets) when it contains duplicates, missing fields, out-of-date entries, conflicting sources, or free-text notes that contradict the structured fields next to them. It matters for AI because AI acts on what it is given without questioning it. A person spots the contradiction and asks. AI produces a confident answer based on whichever version of the data it sees first.
How clean does my data need to be before using AI?
Cleaner than most teams assume, but not perfect. The bar is honest data, not spotless data: you can name the datasets that matter, you know which are reliable and which are not, and you have an owner for the next fix. For the dataset your first AI use case depends on, aim for accurate, de-duplicated and consistent. For the rest, knowing where they stand is enough to start.
Can a small business fix dirty data without a data team?
Yes. Most data cleaning for a smaller business happens in a spreadsheet and a CRM, not a data platform. Export the one dataset that matters, de-duplicate it, complete or flag the critical fields, resolve any conflicts between systems with a written rule, standardise the formats, and name an owner to keep it clean. That sequence is achievable in days, not months, and it does not need specialist tooling.
Carl Chessum is the founder of AI Readiness Partner and the author of AI Readiness for Marketing Leaders, available on Amazon. He has spent 25 years inside transformation programmes, client-side and consultancy-side, across PLC, VC-backed and private-equity-backed businesses.