← Blog/AI Governance·19 March 2025·6 min read

Dark Content Is Quietly Poisoning Your Enterprise AI

TL;DR

  • Your M365 estate is the training set for your enterprise AI.
  • Dark content—outdated templates and clauses—poisons AI output.
  • AI Governance is impossible without content governance at the source.
  • Curate the content AI can access, starting with templates.

A junior lawyer asks your new AI assistant to draft a sales agreement for a US client. The AI, drawing on the firm’s document repository, confidently produces a contract. The partner signing it off does not spot that the AI has used a three-year-old template from a senior associate’s personal OneDrive. The jurisdiction clause is for New York, not the required California, and the liability clauses are long out of date. The error is subtle. The consequences are not.

This is the core challenge of deploying generative AI against enterprise content. The value an AI creates is directly proportional to the quality of the content it can access. Yet for most organisations, the Microsoft 365 estate is a vast, uncurated library of ‘dark content’: a sprawling mass of outdated templates, superseded clauses, and orphaned drafts.

Garbage In, Gospel Out

The principle of ‘Garbage In, Garbage Out’ is as old as computing itself. In the age of large language models (LLMs), the formula has changed. It is now ‘Garbage In, Gospel Out’. An LLM will synthesise erroneous information from a flawed source document and present it with the same confident authority as if it were drawing from a pristine, board-approved board pack.

Internal AI systems, including Microsoft Copilot and bespoke retrieval-augmented generation (RAG) setups, do not possess innate institutional knowledge. They are retrieval and synthesis engines. They find what they believe to be the most relevant information within their designated search space and use it to construct an answer. When that search space—your SharePoint sites, Teams channels, and OneDrive accounts—is littered with duplicates and drafts, the AI is programmed for failure.

The risk is not that the AI will fail spectacularly, but that it will fail subtly. It will insert an outdated figure into a financial report, apply a superseded policy to an HR query, or use a non-compliant footer in a client communication. The output looks correct, feels correct, but is quietly wrong.

The Anatomy of Dark Content

Dark content is more than just old files. It is the digital residue of undocumented workflows and inconsistent information management. It represents a systemic failure to govern the lifecycle of your most critical intellectual property.

In our work with large enterprises, we see this manifest in several forms:

  • Outdated templates saved to local drives, becoming the unofficial starting point for high-stakes documents like MSAs or regulatory disclosures.
  • Superseded legal clauses copied and pasted from old Word documents, creating a trail of contractual risk.
  • Superseded brand assets and disclaimers embedded in PowerPoint and Excel files, leading to inconsistent external communications.
  • Orphaned drafts of financial models in Teams channels, where an AI might retrieve incorrect assumptions for a new forecast.
  • Knowledge-base articles in SharePoint that refer to retired processes, feeding bad advice to AI-powered chatbots.

Your M365 Estate Is the AI’s Training Set

For an AI, your corporate server is not just a filing cabinet. It is a curriculum. Every document, slide, and spreadsheet is a lesson in how your organisation communicates, calculates, and codifies its obligations. Without deliberate curation, you are allowing the AI to learn from a deeply flawed syllabus.

Microsoft’s tools are designed to index this content efficiently. Copilot’s value proposition rests on its ability to ground its outputs in your data. The problem is that it will do so indiscriminately, unable to distinguish between a signature-ready master template and a project manager’s two-year-old draft.

Treating low-level content hygiene as a low-priority IT task is a profound strategic error in the AI era. It is equivalent to allowing your AI to train on random, unvetted data from the internet. The result is the same: unpredictable, untrustworthy, and occasionally hazardous output.

Shrink the Surface, Govern the Source

The solution is not a vast, manual clean-up project. Attempting to review and classify every existing document is an impossible task. The solution is to govern the creation of *new* content at the source and to provide the AI with a curated, trusted knowledge base to draw from.

This means shifting focus from managing documents to managing the templates and components from which they are assembled. By establishing a governed template layer within Microsoft 365, you create a ‘golden thread’ for your content. You ensure that when an employee—or an AI—creates a new document, it starts from a compliant, up-to-date, and approved foundation.

A central clause and asset library, integrated directly into Word, Outlook, and PowerPoint, ensures that boilerplate for legal, finance, and marketing is always the correct version. It removes the need for users to copy and paste from old files, effectively shrinking the surface area from which an AI can pull poisonous content. This is the foundation of a reliable enterprise AI system.

The intelligence of your AI is a mirror, reflecting the quality of the content you feed it. Governing that content is how you ensure the reflection is a true one.

FAQ

Can't we use Microsoft Purview to manage this problem?
Microsoft Purview is essential for applying retention policies and sensitivity labels to existing content at scale. However, it does not govern the creation of new documents. A governed template platform addresses the root cause, ensuring new documents start from a compliant point, rather than retrospectively classifying non-compliant ones.
Is this only a risk for legal and financial documents?
No. The risk is enterprise-wide. An AI can just as easily pull an obsolete brand message into a marketing presentation, an old process into an HR document, or a superseded disclaimer into an Excel model. Any function that relies on standardised documentation is vulnerable to AI error rooted in dark content.
How do we start if our content is already a mess?
You do not need to boil the ocean. Begin by identifying the 20-30 most critical document types—master service agreements, board packs, quarterly reports. Centralise and govern the templates and clauses for these first. By creating a trusted core of content, you can point your AI to this curated source, delivering immediate risk reduction and building from there.
TALK TO US

Ready to see Kameleon live?

Book a 20-minute walkthrough on our Kameleon demo tenant — every feature, end to end.

One governed source. Every document, every channel.

Or email comms@kameleon.app

We reply within one business day.