Concept

Data First, Then AI? It Depends on the Use Case

The common advice to 'clean data before using AI' is often misapplied. While true for 'Signal Engine' use cases like forecasting, it's a value-blocking detour for 'in-process assistance' where AI's strength lies in making sense of existing, scattered, and unstructured data.

When is cleaning data a prerequisite for an AI project, and when is it not?

  • For Signal Engines (yes) — Data cleaning is essential for AI projects aiming to generate insights like forecasts, scores, or "next best actions." These applications require clean, structured, and machine-readable data to produce accurate and valuable outputs.
  • For In-Process Assistance (no) — Data cleaning is generally not a prerequisite for AI projects focused on in-process assistance. For these use cases, AI's value comes from making sense of existing, often scattered and unstructured data to provide immediate support to users.
  • If Data is Already Harmonized (no need) — If data were already perfectly harmonized and structured, the core problem that in-process assistance AI solves (finding information across disparate sources) would largely cease to exist.
  • When Value is Immediate (no prerequisite) — Projects that can deliver immediate value by assisting users with their current tasks, even with current data fragmentation, do not require a prior data cleaning effort.
  • AI strategy
  • data quality
  • AI use cases
  • in-process assistance
  • signal engine
  • data harmonization
  • machine learning
  • data strategy

Auf LinkedIn lesen →

Markdown-Version (zitierfähig)

Data First, Then AI? It Depends Which Use Case You Mean.

Why the data argument only holds for one of two worlds of AI usage.

There is one sentence that comes up in almost every conversation about AI – from consultants, from platform vendors, from a company’s own IT department: “First we need to clean up the data, then we can talk about AI.” It sounds reasonable. It is easy to agree with. And in practice, it blocks precisely those projects that would deliver value the fastest.

The problem is not that the sentence is wrong. The problem is that it is applied across the board – to a use case where it is true, and to one where it means the exact opposite.

Two worlds of AI data usage

At its core, AI does two very different things. Both are called “AI.” Both involve data. But they place opposite demands on data quality.

World 1 – the Signal Engine. This is about using machine learning to derive usable insights from data: a score, a forecast, a next best action, a pattern. These insights are translated into recommendations and connected back into processes. This is the demanding layer – the one that turns raw data into a usable signal in the first place.

And this is exactly where the data argument holds in full: the data has to be clean, structured, and machine-readable. A forecast built on incomplete, contradictory history is not inaccurate – it is worthless. Anyone who starts here without a data foundation is building on sand.

World 2 – in-process assistance. This is not about prediction, but about overview. The user is in the middle of their work – preparing a quote, in a customer conversation, handling a case – and needs the right information, which today is scattered across systems, documents, and inboxes. AI finds it, summarizes it, and makes it usable as a basis for decisions.

And here the exact opposite applies: it is precisely because the data is unstructured and scattered that the value emerges. The disorder is not the obstacle – it is the reason for using AI. If you had to harmonize the data first, you would no longer need the AI at all.

Why the confusion is expensive

Here is what happens in practice: a company wants to support its sales team with AI – for example, so that a sales rep can more quickly access prior terms, technical specifications, and conversation history during the quoting process. A clear World 2 case. Immediately feasible.

Instead, the initiative ends up in a two-year data harmonization project, because the mantra “data first” is reflexively put in front of it. The value is pushed out by years – for a data foundation this use case would never have needed.

The reflex is understandable. Clean data always feels like the right precondition. But for in-process assistance, the migration is not a preliminary project – it is a detour.

The real investment question for CIOs

None of this means a blanket all-clear on data quality – it means a more differentiated way of setting priorities:

For World 2: start now. The scattered, unstructured data is not a blocker – it is the application itself. What matters here is how quickly you can help users in the moment of work, not how clean the data warehouse is.

For World 1: the Signal Engine is the bottleneck – not the AI. The expensive, demanding part is not the AI model on top. It is the layer underneath that turns heterogeneous sources into a reliable signal. Whoever invests here invests correctly. Whoever believes that buying an AI agent settles the matter underestimates where the real work lies.

That is why the decisive question at the start of any AI initiative is not “Is our data ready?” – but rather: “Which world does this use case belong to?” The answer determines whether data quality is the first hurdle or no hurdle at all.

Daniel Gorld

Consulting Director, cbs CX (The cbs Group Salesforce Consultancy)

Daniel Gorld is a B2B CX and process consultant and Consulting Director at cbs CX (The cbs Group Salesforce Consultancy) with over 20 years of experience in industrial B2B.

Mehr über DanielLinkedIn →

Related

Published · Updated