The Data Myth: How Much Data You Actually Need to Get Value from AI
Heads up, this episode opens with a short ad, about 1 minute.
About this episode
Most business owners already know their data is a mess. Spreadsheets in three formats. Customer notes buried in email. Files nobody has opened in years. That mess is not the problem. Waiting for it to disappear before starting an AI project is.
This episode walks through Plan, Do, Check, Act, a framework for scoping only the data one project actually needs instead of trying to fix the whole business first. It includes a real client example built on dark data, Word documents, scattered folders, and years of untouched email, and the two specific ways this goes wrong when owners skip the scoping step or let it expand without limit.
For a small business owner with no data team, this is the difference between a project that launches this quarter and one that stays a cleanup exercise indefinitely.
KEY TAKEAWAYS
- Messy data is normal for every business. The real problem is treating that mess as a reason to delay a project instead of a normal starting condition.
- Waiting for the perfect data set is not caution. It is procrastination wearing a costume.
- PDCA, Plan, Do, Check, Act, works by scoping data to one project at a time, not the whole business.
- A client's project ran on dark data, Word documents, scattered folders, and years of untouched email. Only the slice one AI tool needed got organized. Everything else went on a gap list.
- There are two ways this fails. Skipping the data work and hitting bad results downstream, or chasing every data gap and never launching. The fix for both is the same discipline, scope to the one tool, log everything else, and launch.
- Pick one AI project, map only the data it touches, organize that slice, and resist expanding scope in either direction.
Have a question for Dr. Mike? Visit https://thotosai.com your question may become a future episode.
Take the free AI Readiness Assessment at https://thotosai.com/assessment
Chapters
- 0:00Introduction
- 1:46The Problem
- 2:41The PDCA Framework
- 3:29Dark Data
- 5:15Traps to Avoid
- 6:52Steps and Assignment
Transcript
Waiting for perfect data before you start your first AI project is not caution. It is procrastination wearing a caution costume. Hello, welcome to The Readiness Report. I'm your host, Dr.
Mike Donaldson. Today on the podcast, we're talking about the data myth, the belief that your data must be clean and complete before you start your project. Owners will wait. They will organize their data.
They'll tell themselves, maybe next quarter we'll be ready and we'll get that project going. But that time and that work just continues. The waiting costs is the one thing you cannot get back. The lessons that a real project teaches you, that is where you start making your advancements on your AI journey.
Why you should listen to me. I have 20 years of experience in manufacturing. I've worked on projects from beginning to end. implementing process improvements before machine tools ever came into play or new software.
I have a doctorate's degree in business administration, specializing in entrepreneurship, which is helping businesses grow. I've sat on both sides of the table. I've been the operator and I've been the consultant. And I know exactly what happens when the foundation for any project is skipped and how things can go wrong.
Now, before we start, please subscribe to the podcast. It helps the show out and help other people find me. You can drop a comment or question to thodosai . com and I'll be more than happy to get back to you.
If you would like to know if your business is AI ready, I have a free assessment at thodosai . com that you can take and get instant results. It'll give you a good snapshot of where you are right now. Okay, let's get into the data myth.
So the real problem, it's not messy data. We all have messy data. We've built our files around ourselves. They're designed for people, they're not designed for AI, so all the data is messy.
The real problem is treating that mess as a reason to delay, not as a normal starting condition, which it is for all of us. Cleaning up a project instead of starting, you end up reconciling old records, you're chasing files nobody looks at, nobody has looked at, they've been in archive for years. And it does feel productive. And you do need to clean up your data.
Don't misunderstand me there. But when you're working on a project, you're working on data that the business doesn't need right now. It does not increase your revenue. It does not improve customer satisfaction.
Waiting for the perfect data is not caution. It's procrastination, wearing a costume, and it delays you from getting started. Now, the framework that can help you on this, it's called plan, do, check, act. or PDCA.
It's been around longer than AI, and it's well proven, and it's an iterative process. First, the plan. This is where you scope the data. In this situation, you're scoping the data for that one project, not the whole business.
The do. Now you're going to organize that slice of the data and launch. Check. You're going to look and see what broke.
Something will break. Gaps are going to surface once real use starts with that tool. Now you're going to act. You're going to fix those specific gaps and then decide what is next.
And then you iterate again and you go through plan, do, check, act again. Now, in my example, I had a client. We were working on a project and we're looking at their data. And most of their data is what the industry calls dark data.
What those are, they're Word documents or scattered folders, PDFs, emails that have been archived or sitting that nobody looks at. All of that data is there, but there's no easy way to look at it. AI has made it easier to look at that data. That's why we have to clean it up and start using it.
Now for this project, we did not try to fix all of it. We mapped out only what the tool needed. We created our guardrails. We created our limitations to keep us what was on task and what was outside of the scope of the project.
Anything we found, it went on a separate list or a parking lot. It's a gap list that we would look at later, future projects or future implementations to work on that data. And we did all of the boring work. And trust me, it's boring.
Nobody enjoys pulling data out of PDFs and looking in all of these deep folders and the emails. But this is the foundation work. It is not a detour. It's what's needed.
Then we implemented the project. And of course, there were problems. There always are, but that's the thing. Your project will identify those gaps so you can fix them.
One thing that happened is that all of those were easily resolved. We knew exactly what to do and none of them, none of them came from the data. We had made sure that was clean so there was no rabbit hole for us to fall into and get lost. Our job was to bring that relevant slice of the data into the light and avoid the rabbit holes.
Now, there are ways that this can go wrong, and there's two ways specifically. Not one. Failure one. The team implements a project, but they skip all of that data work.
They simply go to implementation right away and build the AI agent. Now the tool has produced bad results, as expected. But they try to iterate it with the agent. They start tweaking the prompts and adjusting the tool.
That is the rabbit hole. They are going through these iteration steps because they skipped that fundamental foundation work and now everything is showing up downstream. That's not where you want it. You do not want to be dealing with bad data during implementation.
Failure 2. This team, they started working the data. But the scope creep came up behind them. They fixed one gap.
They found another gap. Nearby, so they fix that then they got one more and they started chasing that thread everywhere to every single data problem there was They never launched the project because the definition of ready Kept expanding you need to know what ready means. That's part of the process Now both of these end in the same place no working AI project They didn't get there or they're in an iteration step and they can't fix it or they never started The way to fix both of these is the same discipline.
Scope to the one tool. Establish your guardrails. What is in scope? What is out of scope?
Log everything else and launch on that narrow slice. So this is where you need to start with your project. Step one. Pick one AI project you keep circling back to.
It's the one that you just haven't started. It's on your mind and you haven't gotten there. Step two map only the data the project touches where it lives right now Not the ideal world where it is right now, and that's it This is where you create your boundaries for your project step three organize that one slice If there's anything outside of those boundaries you put in write it down on a list and keep it going That's your parking lot. Don't stop to fix them.
Just write them down and fix only the things that are within your limits Step four. Please don't skip this one, but it does get skipped. Resist expanding the scope in either direction. That includes going down the rabbit holes with the data and getting outside of your boundaries and trying to fix everything.
And then also saying the data is good enough. Make sure you get it done. Don't go halfway. Go the full measure.
You don't want to skip the data work. And don't chase every gap. Those are the things to avoid. Then launch.
Let the tool show you what to fix next. It will point you directly where you need to go, but this time you know it's not going to be your data. This is where you start. Follow PDCA and get that project working as planned.
Now the key takeaway is for you. You don't need all of your data ready. I know the information is out there that you must have your data. Everything's got to be perfect.
It's true. but you need the data for this project. The other things can get fixed later. Your assignment, pick the project you've been putting off, map only the data it touches and get started.
It's much smaller than you think it is. If you get overwhelmed with your data and thinking how long it's gonna take, get past that. The data for your project is much smaller than what you think. Okay.
That's all I've got for you today. But on the next episode, I think we're going to talk about privacy and protecting what you have built before ever opening the door with AI. As always, please take a moment. Give me a review on Apple Podcast or Spotify, wherever you can leave a review.
It'll help me out, helps the show out. You know, it's going to make me feel good and I'll appreciate that. OK, I am Dr. Mike.
Thank you for listening and you all have an excellent day.