The short answer
Do we need clean data before we can use AI?
No, and waiting for clean data is the single most common reason a project never starts. Most useful AI work needs access rather than perfection. The cleaning that genuinely matters happens as part of connecting things, on the data you actually use, which is a much smaller job than cleaning everything first.
Where the myth comes from
The idea is a holdover from an older kind of project, where a company built a warehouse and every report depended on the whole dataset being consistent. In that world data quality was genuinely a prerequisite, and the cleanup came first because nothing worked until it did.
That is not the shape of most AI work now. Answering a question about last month's jobs does not require five years of history to be immaculate. It requires this year's records to be reachable and roughly right, which is usually already true.
What does have to be true
The systems have to be reachable, the important records have to exist somewhere other than a person's memory, and somebody has to be able to say what good looks like. Those are the real prerequisites and they are a much lower bar than clean.
Where data is genuinely bad in a way that matters, the honest answer is to say so and scope the fix rather than build on top of it. That is a specific piece of work on a specific field, not a company-wide cleanup project.
Cleaning as a by-product
Connecting systems surfaces the problems immediately: the duplicate customer, the field three people fill in differently, the records that stop in 2023. That is useful information you did not have before, and it arrives attached to a reason to care.
So the cleanup happens where it pays off, on the data something actually depends on, rather than uniformly across everything on the theory that it might matter one day.
And then
What people ask next.
Who should own the data work on our side?
Usually whoever already knows where things really live, which is rarely the most senior person available. What matters is that somebody can answer questions about how records are actually used and can decide what good looks like for a given field. That is a few hours across a project, not a role.
Does imperfect data mean unreliable answers?
It means the system should say what it is unsure about rather than answer confidently anyway, which is a design decision rather than a data problem. Built properly, it shows where a figure came from and flags what it could not reconcile, so an imperfect record produces a caveat instead of a wrong number.
Access first. The cleaning follows the thing that needs it.
- Also askedWhy do most small business AI projects fail?
- Also askedAI receptionist or answering service: which does a small business need?
- Also askedHow do I choose an AI consultant?
Built in Grand Rapids, Michigan, and put to work wherever your business is.