26-08-2026 Door: Winfried Etzel

The Broken Assumptions of Data Governance

Blog-challenges

Data governance as it is widely practiced today was designed for a world where a human could influence data decisions, be hands-on, and be able to observe, verify, and adjust when needed. That world can no longer be the assumption data governance is built on.

A retail company changes what “active customer” means. E-commerce needs the definition to include anyone who opened the app in the last 90 days to meet their engagement targets.

Marketing uses the definition to flag for win-back offers. Offers for “inactive customers” to come back to the store. Due to the updated definition of “active customer” anyone who stopped buying a while ago, but still opens the app is dropped from the campaign.

Demand forecast adjusts inventory orders to keep up with that growing active customer base. Meanwhile finance keeps to the official definition that reflects regulatory reporting requirements.

As time passes management sees a flat customer base while marketing reports growth and asks for adjusted budgets. First when someone dived into the numbers the different definitions surfaced.

The story is grounded in good intentions and takes structurally correct actions. But meaning changes. If you run a lesson-learned session on incidents like these, the first thing that comes to mind is always communicating. Someone should have seen or said something. We could add checkpoints to the process. We could even automate these checks in the tool.

But it is not that easy. People involved in these examples are competent, the tools work, the processes are followed. The problem is structural, in how we manage and govern data, why we do it, and how we apply frameworks.

We do data governance for how the world was, not how it will be.

We rest on assumptions that are outdated.

Organizational meaning changes. That is semantic drift. And it is more than changes to the technical structure.

The first assumption: humans stay close to the data

Not that long ago, we could solve data governance problems in a meeting, over a cup of coffee, in a workshop. We did not always need a ticket to be raised or the Data Governance Council to approve, or a new policy to be implemented.

Looking back, these were often effective data governance mechanisms available to us. Even if we didn’t call it data governance. Issues could be identified, behavior observed.

Now we live in a time of data abundance, constant reuse, distributed interpretation, and distancing. The data runs through a pipeline, a metrics layer, a forecasting model, a campaign platform, automated based on a repeatable setup.

But most of our data governance mechanisms were designed for a world where a human sat near consequential decisions. This has become increasingly difficult. With AI and AI agents we introduce non-human data actors. The humans who could both identify certain situations and communicate about them are increasingly removed and engineered out of the flow. All while the governance framework is still assuming something else.

Semantic drift is one manifestation of this problem. Humans are no longer close enough to many consequential decisions to continuously observe and amend.

This requires a different kind of governance that is far more embedded in the flow of work.

The second assumption: humans can verify meaning at machine speed

Look again at the example we opened with: active customers excluded from win-back campaigns.

The example encodes a judgment made at a certain point of time with a certain intention. When we think that we automate tasks, we often just freeze a judgment and, worse, scale it beyond its initial intent.

Sometimes our circumstances and context changes, sometimes meaning changes. This requires attention. The extension of this problem is that humans cannot verify at machine speed or scale. Speed and volume are increasing beyond our human capabilities to govern.

Techniques that focus on AI or model governance are finding their way into data governance, because we cannot verify at machine speed. The response cannot be more manual checking. Data governance must move to constraints, feedback loops, and assurance mechanisms as integrated mechanisms in our daily routine.

The third assumption: what works for software, works for data

Software engineering has been a dominant source of methodological inspiration for data for a while. We imported CI/CD to data, introduced agile to data governance, product- and platform thinking, and domain-driven design. But there are some truths to software that do not translate well to data. A service runs where it is deployed. Domain-driven design entails the idea of bounded context. Here, we think of context that we can determine and relate to as relatively stable. Terms are defined as stable within this context.

Yet, data does not respect bounded context. It travels across context, changes meaning, adopts new definitions.

In addition, software is implemented when it is shipped. The data lifecycle needs to be managed way beyond the lifecycle of applications or systems.

Software engineering patterns can enrich our methodologies in data, but they need to be adopted carefully. They help when they make agreements explicit and enforceable. But they cannot treat data as stable, bounded code.

We are preoccupied with whether we could

This is one of my favorite scenes from the movie Jurassic Park. Ian Malcolm, played by Jeff Goldblum, tells the park founder: “Your scientists were so preoccupied with whether or not they could that they didn’t stop to think if they should.”

We focus on all the engineering questions, in control of the how. But we skip the fundamental one: Why?

The way we construct data governance in modern organizations contains the same issue. We focus on the how, and while all our mechanisms of enforcement work, the world they were designed for has changed underneath. We have created data catalogs, quality and observability tools, glossaries, and policy frameworks. An entire industry built on how we do data governance, how we monitor, how we enforce.

We need to remind ourselves why we are governing data in the first place. It is not to implement tools or policies.


Data governance is a human-based system by which data assets in a socio-technical system are directed, overseen, and by which the organization is held accountable for achieving its defined purpose.


Human-based does not mean manually performed by humans. It means accountability, intent and the authority to intervene stays human.

The three broken assumptions are what led me to think first principles: why are we governing data? The definition above is my answer. We govern to achieve our purpose, through negotiating data requirements, directing data towards a goal, and assuring that we can be accountable along the way.

I call this the Data Governance Triad: negotiation to make meaning explicit, direction to bind us to our deliberate choices, and audit to ensure accountability over time.

Once we peel away years of how-to, of enforcement and monitoring practices, and think of the core of what data governance is, it becomes easier to change our approach and tackle the challenges ahead.

The Triad as a way to focus on the purpose of data governance, is at the core of my book Data Governance in the Wild.

The question that remains

AI did not create these governance issues but has exposed them. Machine reasoning now operates on our data at scale, drawing conclusions, triggering actions, making decisions. All with full confidence.

We have asked and explored what machines can and cannot do. But the more important question is what a machine should or should not own. Who stays accountable for machine reasoning? Who answers for a decision where no human was involved?

The answer to this question is to be found in the socio-technical reality we operate within, in treating data governance as a capability the organization performs. This is where Data Governance in the Wild places data governance: at the core of the socio-technical system, between humans and machines.

It starts with negotiating meaning between two teams that hold different definitions, directing them to an explicit choice, and assuring that their job can run at machine speed.

The practical starting point is to identify critical terms, decisions, and data products where meaning crosses domains. This is where you need to make agreement explicit, encode what can be encoded and create an assurance loop that detects when the agreement no longer holds.

This new reality requires us to rethink and to find new ways of embedding data governance in how we govern organizations at large. We have governed data in a way that no longer is suited for our reality.

Geef een reactie

Je e-mailadres wordt niet gepubliceerd. Vereiste velden zijn gemarkeerd met *

Adept Events

  • thumb image

    Winfried Etzel

    Winfried Adalbert Etzel heeft bijna vijftien jaar ervaring in data- en informatiemanagement, met een focus op data governance, metadatamanagement en data- en AI-strategie. Hij is bestuurslid van DAMA Norway en actief binnen de Scandinavische datagemeenschap. Hij is verbonden aan Data Management Advisors, een geregistreerde opleidingsaanbieder voor DAMA-International en de enige gecertificeerde opleidingsaanbieder in West-Europa.
    Daarnaast is hij host van de podcast #MetaDAMA, waarmee hij data-professionals in de Scandinavische landen een stem geeft en een holistisch beeld schetst van datamanagement. In zijn visie moeten digitalisering en data-analyse aansluiten op de behoeften van zowel de organisatie als haar klanten, zodat organisaties meer waarde uit hun data en informatie kunnen halen. Winfried spreekt ook voor Adept Events en treedt regelmatig op tijdens de DW&BI Summit in Utrecht.
    Alle blogs van deze auteur
Deze website gebruikt cookies om de beste gebruikerservaring mogelijk te maken. Meer informatie