All episodes
    Season 1, Episode 99 min

    Privacy and AI: How to Protect What You Have Built Before You Open the Door

    Heads up, this episode opens with a short ad, about 1 minute.

    About this episode

    DESCRIPTION

    You think you know your business data. Names, emails, purchase history, all accounted for. Then you actually go look, and the picture changes. Records live in a CRM. More sit in email threads nobody archived. PDFs sit in a shared folder nobody has opened in years, and phone numbers get duplicated across systems that never talk to each other.

    This episode walks through why privacy is a readiness problem, not a legal one, and why it has to be settled before any AI project starts. Dr. Mike breaks down what dark data actually is, walks through a real client example where a security gap surfaced before AI ever entered the conversation, and lays out the steps for building a data inventory you can actually use to decide what AI is allowed to touch.

    A ten-person business is just as attractive a target as a much larger one, and most small business owners have never stopped to ask where their own data actually lives.

    KEY TAKEAWAYS

    Dark data is data that exists, you know it exists, but it is in a format that cannot be searched easily. It is not clutter, it is exposure.

    The digital twin pillar of AI readiness means building an accurate picture of what data exists and where it lives, and that picture is the gate nothing moves through until you know what is on the other side.

    A real service business found a backdoor in its web-based CRM that let outside parties download full client records, and it was only found because the business stopped and asked how its data was actually protected right now.

    Closing a security gap protects against two threats at once: outside intrusion and unauthorized movement of data by AI. Same gap, two different threats, close it once.

    A ten-person business is just as attractive to someone after data as a much larger company, sometimes more so, because the guardrails were never built in the first place.

    Where to start: build a data inventory that includes your dark data, and label each area usable, off-limits, or read-only and contained.

    Have a question for Dr. Mike? Send it to mdonaldson@thotosai.com or visit thotosai.com — your question may become a future episode. You can also take the free AI Readiness assessment to see where your business scores https://thotosai.com/assessment

    Chapters

    1. 0:00Introduction
    2. 1:36Dark Data
    3. 2:58The Digital Twin Framework
    4. 4:31Traps to Avoid
    5. 5:35Steps to Get Started
    6. 8:02Assignment and Close

    Transcript

    If you cannot find it, you cannot protect it. And if you cannot protect it, you cannot safely hand it to an AI tool, no matter how good that tool is. Hello, welcome to the podcast. I'm your host, Dr.

    Mike Donaldson. Today, we are talking about privacy and AI and whether you know where your sensitive data truly lives. If you do not know what data you have, you cannot decide what AI is allowed to touch. Privacy is not a legal problem.

    This is an AI readiness problem and it comes before AI. You have to know where your data is to ensure that it's safe and protected from AI doing anything with it that you do not want, anything silly. I have 20 years of experience in lean and manufacturing. I have a doctorate's degree in business administration and entrepreneurship and I have sat on both sides of the table.

    I have been there as an operator and as a consultant and I know firsthand what can happen if you do not have the foundation in place before you purchase a machine, a software tool, and an AI tool. Before we get into it today, please take a moment and subscribe to the podcast. It's easy You can just click a button and if you have any comments or if you'd like to know about Your company and where you stand in AI readiness at thotusai . com I created a free readiness assessment It takes about 10 minutes and you can get a quick snapshot and some ideas of where to start Every business thinks they know their data the names the emails the purchase history.

    It's all there. We know it's there Then when you start diving into it, the real picture forms. The data is scattered around. The records can be in a CRM.

    They could be in an archived email thread that no one's organized. They could be in PDFs that nobody has opened in years in a shared folder that also has never been touched for years. Phone numbers can be duplicated across several different systems. This data is called dark data.

    What that means is it's data that exists, you know it exists, but it's in a format that cannot be searched easily. The PDFs, they're not easily searched, etc. Dark data, it's not clutter, it's exposure. If you cannot find it, how are you going to protect it?

    You can't safely hand that to AI without knowing where it is and how you have it protected. In the shop floor, parts can go missing all the time. But nobody notices it until there's a count, either an inventory count or during the build a part is missing or after the build parts are missing. All of those are the same thing that can happen with AI.

    The digital twin pillar of AI readiness is where all of this fits in. You can create an accurate picture of what data exists and where it lives, and that is the digital twin. That picture is also the gate. Nothing moves through until you know exactly what is on the other side.

    That is how you protect it. I worked with a client in the service business. They had clients coming in regularly. They had a web -based CRM.

    We were looking at doing an implementation, and we started looking at the data first. We started asking hard questions about security and then a gap surfaced. There was a back door within the website that no one had noticed. Three, two, one.

    The back door let outside parties in and they could download and view full client records. They had access to the contact info, appointment history, purchase history, everything. The only reason this was found was because the business stopped and asked, how is our data actually protected? right now, not how we think it is.

    We need to go look and find out. Now once you close that gap, you protect it on both outside intrusions and unauthorized movement of the data by AI. You have the same gap with two different threats. Close it once and you can protect against both.

    Know where your data is, know how exposed it is, and fix it. Only then, only then, decide what AI can touch. Where this can go wrong is owners can treat privacy as a compliance check the box it can be handled before AI but not in the detail or It's done after the AI tool is already running by this point. The tool has already touched the data Another problem that occurs is Businesses will think we're too small.

    They're not going to come after our data. Why would they there's nothing there? but a 10 -person business is just as attractive to someone that wants to get to the data. You're not secure, and oftentimes, your business may not have the guardrails in place to protect it or the security measures.

    Then there's a quiet mistake. There's no rules in place at all for AI. You've built some AI pilots. You put no guardrails in place or no rules, and that AI could be doing things with your data that you don't know.

    You could have your business locked down and protected from all outside intruders and still leave a door inside for the AI to copy your data right out your door. Security and AI governance are related, but they are not the same guardrail. You need to do both. You need to have the rules in place and the guidance for the governance, and you need to have the security measures in place to stop the AI from doing anything you don't want.

    Now, where do you start? And I'm going to be clear on this. There's four steps, but this is what is not in the assignment this week. You are not doing a full data cleanup.

    Okay, let's go to step one. Step one, build a data inventory that includes all of your dark data. We're talking about the CRM, emails, PDFs, spreadsheets, shared drives, every place your data lives. Label, each area, three labels.

    One is usable. These are the files. that AI can use for their project. The files that are off limits, that's your second label.

    These are the files that AI is not to touch, not even to look at. It can be entire folder directory, but these are off limits. The other is a read only and contained. What that means is the data is viewable and nothing can get copied out.

    It can only read it, it cannot copy. Nothing leaves the network. You can adjust that one as well. Maybe you need to have files that you want AI to only read and that's fine.

    They can copy and take what they need. But this is where you have your guardrail someplace as well. But if you only want them to read data just as a reference, so it's a knowledge base, then that's read only and contained. They can do nothing with it.

    Step three. Look for the same kind of gap that I talked about earlier. Ask directly if someone wanted in and where they would go to get in and how they would do it. Your team and yourself, you may know of a way that someone could get in if they wanted to.

    Those are the gaps in the doors that you need to close in lockdown. Step four, and this one, people can skip this or it can be overdone. Only clean the data your AI needs first for the AI project. Nothing else.

    You only do... what the AI needs and what it needs to touch. This goes back to episode eight on how much data you actually need before you start. You are cleaning what is in the path for that project that you are doing right now.

    Nothing else. Now your key takeaways from today I can put into one sentence. You cannot protect data you have not found and you cannot safely use AI on any data you have not decided how to handle. Your assignment for this week is build that data inventory, including your dark data, label each one of those areas as usable, off -limits, read -only, and contained.

    That's how you start to protect your privacy. And then that information, you can use that for any AI project you're doing in the future. That's all I have for you in this episode. In our next episode, we will talk about the human element and what AI cannot replace and why that is good news for all of us.

    Please take a moment, provide a review for this podcast. It will help the show and it will also help other people find the show just like you did. And if you have received anything useful from anything we have talked about in the past, please share this episode with someone else so that they can get exposed to it as well. I'm Dr.

    Mike. Thank you for listening and you all have an excellent day.