Information Hygiene: The Security Discipline That Matters More Than Ever

For twenty years we secured the perimeter. Firewalls, endpoints, zero trust. All of it built on one idea: danger comes from outside, and we can stop it before it reaches people. That idea still holds. It's just not the whole story anymore.
What you collect, where you share it, who can reach it, when you delete it. That's a risk now too. And AI didn't just add to the problem. It made the problem cheap to exploit.
IBM titled its 2025 Cost of a Data Breach Report "The AI Oversight Gap" because that's what happened: companies adopted AI faster than they learned to control it. This post argues for information hygiene as its own discipline, an operations one, not a developer one. It happens where work happens: SharePoint, OneDrive, Drive, Dropbox, Teams, and the folders nobody remembers making.
What information hygiene means
Take care of your information from start to end. Collect only what you need. Label what you have: you can't protect or delete what you can't find. Give people the least access that works, because a file open to everyone is only as safe as the weakest password in the company. Delete what you don't need; extra data is cost, exposure, and risk all at once (TechTarget). And check where a piece of information came from before you trust it.
None of this is new. What's new is what happens when you get it wrong. An old SharePoint site full of contracts and HR files used to be somebody's mess. Now it's fuel.
Your collaboration tools are already too open
Before we get to AI, look at where your files actually sit.
Varonis scanned SaaS tools in 2024 and found one in ten sensitive files open to every user. Microsoft 365 made up nearly half of that. Average risk per organization: $28 million. A related study found 53% of companies have 1,000+ files available to every single employee.
The worst cases are the forgotten ones. Microsoft's own reviews flag inactive SharePoint sites as the highest-risk group: the site built for a project that ended in 2022, still fully open, owned by someone who left.
Google Drive has the "anyone with the link" problem. One click during a two-minute call puts a document on the open internet. Metomic scanned about 6.5 million Drive files: 40% held sensitive information, a third were shared outside the company, and half a percent were fully public, almost all by accident. And it doesn't fix itself. Japanese developer Ateam left personal data of nearly a million people open through a Drive mistake for over six years.
Dropbox isn't worse than the rest. But in April 2024 its Dropbox Sign service was broken into: emails, phone numbers, login data, disclosed in an SEC filing. One company was hit; every customer that used it for signing was exposed with it.
The tools themselves get used against you, too. People are about seven times more likely to click a bad SharePoint or OneDrive link than a normal phishing link, because it's a link to something they trust.
That was the starting point before AI arrived. Messy permissions, dead sites, public links, forgotten app grants. AI raised the stakes.
Three ways AI changes the math
It takes your files with it
Anything an employee pastes into a chatbot, or a "connect your data" feature reads, is out of your hands. Cyberhaven's 2026 report found the average employee puts sensitive data into an AI tool about once every three working days. Two-thirds of staff at large companies have used AI tools nobody approved. IBM found unapproved AI behind 43% of AI-related breaches, double the year before, and one in five companies had a breach tied to it outright.
Remember those one-in-ten open files? That's what an AI assistant reads the moment you connect it. Obsidian researchers note that AI tools, once given access, can scan and copy data as it changes. The biggest SaaS break-in of 2025 started with stolen login tokens, ten times the usual damage. And a flaw in Microsoft's own OneDrive File Picker gave apps access to far more than the files people picked, silently, to millions.
One thing worth knowing: "not used for training" does not mean "not stored." Two different promises. Ask how long inputs are kept and who can read them meanwhile. The serious vendors are clear. Google promises not to train on your data without permission, for instance, and offers zero-retention setups.
IBM also found breach costs fell for the first time in five years, to $4.44 million, thanks to AI-assisted detection. AI is on both sides. Which is exactly why hygiene, not avoidance, is the right answer.
It makes attackers better at pretending to be people you know
Voice phishing rose 442% between 2023 and 2024. Deepfakes grew 680%. Credential phishing jumped 703% in the second half of 2024 alone (Brightside AI). The famous case: a fake CFO on a video call talked a finance employee into sending $25 million.
What made that attack work wasn't the fake video. It was everything behind it: the public org chart, the conference talk, the internal files a stolen account could read. Attackers can learn about you without breaking in. Publish less about your structure, your tools, your people's schedules. That's half the defense right there.
It dirties your information from the inside
Marketing pastes AI output into the SharePoint wiki. Analysts save chatbot summaries into Shared Drives. Slowly the question changes from "is this email really from my CFO" to "was this ever true." Models trained on their own output get worse over time. Researchers showed this in Nature, and IBM describes the same effect. Unmarked AI text that lands in your wiki today is the "truth" your team works from next year.
One more thing: cloud drives are not safe from ransomware. Proofpoint showed an attack that locks SharePoint and OneDrive files for good by burning through version limits (OneDrive's default is 500), leaving nothing but locked copies and a ransom note (Dark Reading). Ransomware on a synced laptop travels up, because encryption looks like any other edit (Spin.AI). A tidy library is cheaper to recover than a pile.
What to actually do
Count your files so you can delete them
Go through SharePoint, OneDrive, Drive, Dropbox, Teams, and file servers with deletion as the goal. Modern tools can scan at scale and flag what should be kept or deleted (CIO Influence). Label data by how sensitive it is and what AI may touch it: fine for outside tools, fine for internal assistants, or never. If your labels say nothing about AI, they're out of date.
Check who can open what
For every SharePoint site and Shared Drive: who owns it, who can read it, does it still need to exist. Archive the dead ones. Turn off "anyone with the link" and give people named, safe sharing options instead, so the easy way is the safe way (Material Security). And remember: an AI assistant gets the same access the moment it connects. Wiki open to the whole company? Your assistant will summarize it for the whole company.
Review every app you've connected
List every outside app with access to your storage. Does its access match its job? Revoke the rest. Dropbox Sign and the OneDrive File Picker teach the same lesson: the file you didn't pick is the one that hurts you.
Write an AI policy people can follow
Banning AI doesn't stop people using it; it stops you seeing them do it. They just move to personal accounts with no rules (Adaptive Security). If sharing a file the approved way is painful, people use personal Dropbox. Same instinct. Pick work tools with real data terms, ban personal AI for sensitive work, write a simple policy (Axiom). Then make the approved way the easiest way.
Decide how long you keep things
About 60% of companies using AI have no rules for how long they keep the data involved (VerifyWise). Keep too long and risk grows; delete too soon and you lose your legal defense. Write a schedule, and make it agree with your records rules: privacy and records policies clash more often than not (IAPP).
Check big requests twice
Payment orders, password resets, urgent asks from the boss. The $25 million deepfake would have died on one phone call back. And mark AI-written material before it lands in your knowledge bases.
Watch what you publish
Your website, LinkedIn posts, conference talks. What do they teach an attacker about your structure and your people's schedules? Publishing less is hygiene, not paranoia.
Train for the good fakes
Not "don't click links" but "expect perfect grammar, correct context, your colleague's real voice, and slow down when it's urgent." Correcting someone the moment they share something risky beats the yearly training video (Adaptive Security).
The uncomfortable conclusion
Information hygiene is boring. Nobody demos a permission review at a conference. But AI flipped the deal: attacking messy information got cheap, and cleaning it up didn't get more expensive. IBM's data shows who pays: companies that adopted AI without rules took the losses, and 85% of breached companies plan to spend more on governance after the fact.
The perimeter won't save you from an AI assistant that already has legitimate access to your SharePoint or Google Drive. The only real defense is information that's clean where it's used: small, labeled, contained, checkable, deleted when it stops being useful.
Hygiene, not heroics. The adventurous cybersecurity life belongs in movies, not reality. Careful and meticulous work is what this moment rewards.
Sources
- IBM Cost of a Data Breach Report 2025, 2025
- Cybersecurity Dive, breach costs and AI governance, 2026
- Varonis, The Great SaaS Data Exposure (2024), 2024
- Varonis, limiting company-wide exposure
- AvePoint, data overexposure causes and fixes, 2026
- Material Security, human error in Google Drive (citing Metomic)
- Valence Security, Google Drive mistake (Ateam case)
- Proofpoint, SharePoint/OneDrive phishing and account takeover
- Proofpoint, cloud ransomware attack in Microsoft 365, 2022
- Dark Reading, Office 365 files open to ransomware, 2022
- Spin.AI, OneDrive ransomware protection guide
- Dropbox, SEC filing on the Dropbox Sign incident, 2024
- Obsidian Security, OneDrive third-party app and AI access risks, 2025
- Infosecurity Magazine, OneDrive File Picker flaw, 2024
- Adaptive Security, shadow AI risks (citing Cyberhaven 2026), 2026
- Questa AI, shadow AI numbers (citing PagerDuty, IBM), 2026
- Brightside AI, AI spear phishing numbers, 2026
- StationX, phishing numbers (the $25M deepfake CFO case), 2026
- Right-Hand, deepfake voice attacks (citing Google Cloud), 2025
- Nature, AI models collapse when trained on their own output, 2024
- IBM, model collapse explained
- TechTarget, CISO's guide to data minimization
- CIO Influence, data minimization strategy, 2024
- Axiom, employee AI privacy risks for legal teams
- NHIMG, AI data retention policy glossary
- VerifyWise, data retention policies for AI
- IAPP, creating a data retention policy for AI
- Google Cloud, Gemini zero data retention documentation
