← Blog/Governance·5 September 2024·7 min read

Metadata vs Folder Structure: The Argument Enterprises Keep Losing

TL;DR

  • Folders are a failed abstraction for enterprise content; a document has many contexts, not one home.
  • Relying on folder paths for retention and discovery is a high-risk compliance strategy.
  • Capture metadata at creation, inside the template, to make it painless and accurate.
  • Key metadata: Entity, Jurisdiction, Confidentiality, Retention Class, Document Type, and Matter ID.

A General Counsel needs to find all master services agreements with a specific counterparty, governed by German law, signed since 2022. The corporate file share, a labyrinth of departmentally-named folders and cryptic sub-folders, returns thousands of results. Some are drafts, some are duplicates, and the one signed original is nowhere to be found. After two days, the legal team gives up.

This scene is not a failure of search technology. It is a failure of information architecture. The belief that millions of documents, created by thousands of people, can be organised using a folder structure is an abstraction that has long outlived its utility.

The Filing Cabinet in the Cloud

The folder-and-subfolder model is a digital metaphor for a physical filing cabinet. It was designed in the 1970s for single-user systems. It assumes a document has one logical place, a single “home” where it belongs. This assumption is fundamentally incorrect in a modern enterprise.

A single Master Services Agreement, for example, is not just a legal document. It is a sales document, a finance document, and a compliance artefact. To the legal team, it belongs in a folder for client agreements. To finance, it belongs with documents related to that entity’s billing. To the account manager, it belongs in their client folder. The result is either that the document is saved in one place and invisible to others, or, more likely, it is saved in three different places, leading to version chaos.

Findability and Productivity Collapse at Scale

When an organisation has a few thousand documents, a disciplined folder structure can work. When it has tens of millions, that system collapses into a state of high entropy. No single person understands the complete folder hierarchy, and different teams inevitably develop their own conflicting conventions.

This is not a trivial productivity drain. Knowledge workers spend a measurable part of their week—often two to four hours—simply looking for the information they need to do their jobs. They are not searching for new knowledge, but re-locating existing corporate assets. When they cannot find the canonical document, they often create a new one, exacerbating the problem of content sprawl and duplication.

The High Price of Non-Compliance

Basing information governance on folder location is a high-risk strategy. A folder innocuously named “Project Phoenix Archive” provides no machine-readable signal about its contents. Are these documents subject to a seven-year retention policy, or are they under a legal hold? Can you responsibly act on a GDPR Article 30 “right to be forgotten” request when your primary retrieval mechanism is a folder tree?

Regulators and courts are not interested in the nuances of your shared drive’s structure. They expect you to be able to place a hold on, retrieve, and dispose of content based on its substance. Folder paths are merely location data; they are not reliable evidence of a document’s type, confidentiality, or retention status.

A Better Way: Metadata at the Point of Creation

Instead of relying on *where* a document is saved, a modern approach relies on *what* a document is. This is the role of metadata: structured data about the document that travels with it, independent of its location.

Forcing users to fill out a complex form on every save is a recipe for failure. People will choose the first option in a dropdown or enter nonsense just to close the dialog box. The key is to capture most of the metadata automatically at the point of creation, within the document template itself.

When a user opens a

Making Metadata Stick

A governed template layer is the most effective way to solve this. When a user in Word opens a template for a Statement of Work, the system already knows the Document Type. Based on their user profile, it knows their department and office, which can infer Jurisdiction. All the user needs to provide is the client name or a Matter ID.

The interface should prompt for this information in context, within the application where the user is working. It should present managed lists of client entities and project codes, not free-text fields. This eliminates typos and ensures the metadata is consistent and therefore usable for search, filtering, and retention rule processing.

This approach transforms metadata from a chore into a largely invisible, automated background process. Staff are not burdened, and the organisation gains a rich, structured dataset that allows it to manage its document assets with precision.

Which Metadata Fields Actually Matter?

A common mistake is to start with an exhaustive, academic list of dozens of metadata fields. This creates unnecessary friction. The goal is to be effective, not exhaustive. For most enterprises, a small set of core fields provides 80% of the value for retrieval, access control, and retention.

Start with these, ensuring each is a choice from a managed list, not a free-text field, to maintain data integrity.

  • Entity / Counterparty Name: The organisation the document relates to.
  • Jurisdiction: The applicable legal region, which often governs retention.
  • Confidentiality: A simple classification like Public, Internal, or Confidential, to drive access rules.
  • Document Type: The nature of the document, e.g., MSA, Board Pack, SOW, Press Release.
  • Retention Class: The business or legal rule defining its lifecycle, e.g., 7 Years, Permanent, Defensible Deletion.
  • Matter/Deal ID: A project code that links the document to a wider business context.

The argument between folders and metadata is not a technical preference. It is a strategic choice. An organisation that continues to rely on folder hierarchies is choosing to manage its information as a disorganized, high-risk liability. Adopting a metadata-first model is the only scalable way to treat documents as the governed, retrievable assets they are supposed to be.

FAQ

Isn't asking for metadata just more work for users?
No, if implemented correctly. A modern system automates most metadata capture from the user's context and the template they choose. It should prompt only for a couple of critical fields, which is less effort than navigating a complex folder tree to save a file.
What about SharePoint and Teams? Don't they use folders?
They surface a folder-like view, but their underlying power comes from metadata columns and content types. The problem is that few organisations configure this correctly, leaving users with a default, ungoverned experience that mimics a basic file share, defeating the platform's purpose.
Can we apply metadata to our millions of existing documents?
While AI tools can assist with retrospective classification, it is an expensive and imperfect process. The most pragmatic approach is to enforce metadata capture for all new documents from this point forward, and then to triage legacy content based on risk and business value.
TALK TO US

Ready to see Kameleon live?

Book a 20-minute walkthrough on our Kameleon demo tenant — every feature, end to end.

One governed source. Every document, every channel.

Or email comms@kameleon.app

We reply within one business day.