What data readiness checks should you run before adopting AI?
Five pass/fail data questions to run before you buy any AI tool, each with a test you can complete this month and a fix if it fails.
By Carl Chessum
Five checks, and you can run all five inside a month. Can you produce a clean, accurate list of your best 500 customers today? Is the data behind your main decisions complete enough to act on? Do your key systems exchange data without somebody manually exporting a spreadsheet? When an important decision gets made, does reliable data actually drive it? And can you name the person responsible for data accuracy, by name, without checking an org chart? Each one has a pass/fail test below. Fail two or more and the AI tool you are about to buy will not work, no matter what the demo showed you.
This is not a moral position about tidiness. Gartner expects organisations to abandon 60 percent of AI projects that lack AI-ready data through 2026. In Informatica’s CDO Insights 2025 survey of 600 global data leaders, 43 percent named data quality, completeness and readiness as the top obstacle stopping AI initiatives reaching production. Those two numbers describe the same failure from two angles: the model was fine, the inputs were not.
Why these five and not fifteen
There are longer data-maturity frameworks. Most of them are built for organisations with a data function, a catalogue, a lineage tool and somebody whose whole job is governance. If you have 200 people and one very tired IT manager, a 47-point maturity matrix is a way of not starting.
These five are the questions I have watched decide whether a project survives contact with reality. They map to the Data pillar of the AI readiness audit, which is 30 questions across six pillars, takes 7 minutes, and gives you a score out of 120 with a band per pillar and a short synopsis. No sales call unless you ask for one. What follows is the same five Data questions, unscored, with the tests I would use if I were sitting in your office.
Run them in order. The first one is the fastest and it tells you the most.
Check 1: the 500-customer test
The question. If you needed a clean, accurate list of your best 500 customers right now, how confident are you in what you would get?
The test. Ask whoever owns your CRM for it. Do not explain why. Start a timer. Note two things: how long the list takes to arrive, and how much manual work went into producing it.
Pass. It arrives inside two working days, from one system, with no spreadsheet surgery, and the person who received it does not need to caveat any of the fields.
Fail. It takes a week. Or it arrives as a merge of three exports. Or the answer includes the phrase “depends what you mean by best”. That last one is the interesting failure, because it is not a data problem, it is a definition problem wearing a data problem’s coat.
What to do this month. Write down what “best customer” means in a single sentence with a number in it. Revenue over the last 12 months. Gross margin. Repeat order count. Pick one, agree it with your finance lead, and put it in writing. Then have someone build that list once, properly, and record every manual step they took. That list of steps is your remediation backlog and it will be shorter than you fear.
Check 2: how complete is your decision data
The question. How complete is the data your business relies on to make decisions?
The test. Take the last board pack or management report. Pick the three numbers that led to an actual decision. For each, trace it back to source and ask: what percentage of records feeding this are missing a field that matters? Not every field. The ones that matter.
Pass. You can trace all three to a named system, and the gaps are known, quantified and small enough that nobody argues about them in the meeting.
Fail. One of the three is produced by a person rather than a system, and nobody knows the fill rate. Missing data is survivable when you know how much is missing. It is dangerous when the gaps are invisible, because AI tools do not flag a sparse field. They average across it and hand you a confident answer.
What to do this month. Run a fill-rate count on the five fields that matter most in your primary system. Industry, region, owner, product category, whatever your business runs on. One SQL query or one pivot table. Put the percentages on a single page. If a field is under 80 percent complete, it cannot be a filter in anything automated until it is fixed.
Check 3: can your systems talk without a human in the middle
The question. Can your key business systems share data with each other without someone manually exporting and importing spreadsheets?
The test. Draw your five most important systems on a piece of paper. CRM, finance, ERP, support desk, whatever yours are. Draw a line for every data flow between them. Then mark each line: automated, or a person. Count the people.
Pass. Zero human-mediated flows between the systems that feed your reporting. Or one, and it is documented, scheduled and owned.
Fail. Two or more. This is the most common fail in businesses between 50 and 1,000 people, and it is the one that most reliably kills automation projects, because every manual export is a point where the data stops being current and starts being a snapshot somebody made on a Tuesday.
What to do this month. Pick the single flow that touches the most decisions and cost it out to automate. Not all of them. One. Get a number from whoever maintains the systems. Compare that number to the licence cost of the AI tool you were about to buy. The comparison is usually instructive.
If you have no internal technical capacity for this, that constraint changes what you should attempt first rather than ruling it out. I have written separately about what a small business can actually do without a tech team.
Check 4: does reliable data actually drive decisions
The question. When an important business decision needs to be made, how often does reliable data actually drive it?
The test. List the last five significant decisions your leadership team made. Pricing, hiring, a market you entered or left, a supplier you dropped. For each, name the data that drove it and where it came from. If the honest answer is experience and judgement, write that.
Pass. Three or more of the five trace to data that existed before the decision, not analysis produced afterwards to justify it.
Fail. Fewer than three. This is the check that offends people, so let me be plain: experience-led decisions are not wrong. Businesses have been run well on judgement for centuries. But a business that does not currently use data to decide will not suddenly start using it because a model is now producing the numbers. The habit has to exist first. AI amplifies a decision culture. It does not install one.
What to do this month. Take one recurring decision, ideally a monthly one, and require a named data input before it is made. One decision, one input, one month. If the leadership team ignores the input, you have learned something more valuable than any tool would have told you.
Check 5: name the person who owns data accuracy
The question. Who in your business is responsible for the quality and accuracy of your data?
The test. Say a name out loud. Now ask that person whether they know it is their job.
Pass. The same name comes from you, from your IT lead and from that person. Three matching answers.
Fail. The answer is “IT”, or “everyone”, or a department. “Everyone” means nobody, and departments do not get performance reviews. This same accountability vacuum shows up one level up, in AI spend itself: Open Future Forum’s August 2026 data found business-unit sign-off on AI purchases fell from 18 percent to 8 percent inside a single month while “no single owner yet” doubled to 14 percent. Notably, not one CEO or founder respondent reported an unowned AI purchase at their company. One level down, 29 percent of technology respondents did. The people at the top do not know it is happening. If you have not yet settled who owns AI outcomes in your company, the data ownership question is where that argument starts.
What to do this month. Put a name against data accuracy for your two most important systems. Not a committee. A person, in writing, with the specific fields they are accountable for. It can be a partial role. It cannot be an implied one.
Reading your five results
Five passes. Genuinely rare below 1,000 employees. Your data is not the thing blocking you. Look at process next, because the second most common failure I see is clean data feeding an undocumented workflow that three people run three different ways.
Three or four passes. You can start something narrow. One use case, one data source you trust, one owner. Do not let a vendor talk you into a platform-wide deployment on the strength of the checks you passed.
Two or fewer. Do not buy the tool yet. Not never, yet. Three-quarters of CIOs surveyed by Dataiku said they had remorse over at least one major AI vendor or platform selection made in the past 18 months. Those purchases were not made by careless people. They were made under time pressure, on the assumption that the data situation would be sorted out in parallel. It rarely is, because remediation has no launch date and no press release.
The sentence to watch for
“We’ll fix the data once the AI is in.”
Somebody will say it, and it will sound pragmatic. It is the most expensive sentence in the room, because it inverts the dependency. The tool is bought against a business case that assumes clean inputs, the inputs stay dirty, the outputs are wrong, and the tool gets blamed for a problem it inherited. RAND research published in 2025 found 80.3 percent of enterprise AI projects fail to deliver promised business value, with 33.8 percent abandoned before production and 28.4 percent reaching production but missing expected value. Nobody in that 28.4 percent set out to build something that half-works.
The five checks above cost you a month of somebody’s part-time attention. Running them before the purchase order is signed also gives you something useful when your CFO asks what the money is for: a specific list of what you fixed, and why the case now holds. That is a better conversation than the one where you explain why the pilot stalled. I have written about the cheapest data mistake to fix first if you want to go deeper on remediation sequencing.
Start with the 500-customer list. You can have the answer by Thursday.