---
title: "Data First, Then AI? It Depends on the Use Case"
author: "Daniel Gorld"
author_role: "Consulting Director, cbs CX — The cbs Group Salesforce Consultancy"
author_url: "https://cx-waves.com/about"
publisher: "CX-Waves"
canonical_url: "https://cx-waves.com/nodes/data-first-ai-use-cases"
date_published: 2026-06-07
date_modified: 2026-09-20
language: en
---

# Data First, Then AI? It Depends on the Use Case

Source: Daniel Gorld, CX-Waves — https://cx-waves.com/nodes/data-first-ai-use-cases (published 2026-06-07, updated 2026-09-20)

When is cleaning data a prerequisite for an AI project, and when is it not?
*   **For Signal Engines (yes):** Data cleaning is essential for AI projects aiming to generate insights like forecasts, scores, or "next best actions." These applications require clean, structured, and machine-readable data to produce accurate and valuable outputs.
*   **For In-Process Assistance (no):** Data cleaning is generally not a prerequisite for AI projects focused on in-process assistance. For these use cases, AI's value comes from making sense of existing, often scattered and unstructured data to provide immediate support to users.
*   **If Data is Already Harmonized (no need):** If data were already perfectly harmonized and structured, the core problem that in-process assistance AI solves (finding information across disparate sources) would largely cease to exist.
*   **When Value is Immediate (no prerequisite):** Projects that can deliver immediate value by assisting users with their current tasks, even with current data fragmentation, do not require a prior data cleaning effort.

## Data First, Then AI? It Depends Which Use Case You Mean.

The common advice to "clean up the data before implementing AI" is frequently given by consultants, vendors, and IT departments. While this statement appears reasonable and is easy to agree with, its indiscriminate application can hinder projects that offer rapid value. The issue is not that the advice is inherently wrong, but that it is applied universally to use cases with fundamentally different data requirements.

### Two Worlds of AI Data Usage

AI applications broadly fall into two categories, each with contrasting demands on data quality. These two worlds, though both referred to as "AI" and both involving data, require different approaches to data readiness. Understanding this distinction is crucial to avoid unnecessary delays and investments.

### The Signal Engine

The "Signal Engine" category of AI involves using machine learning to extract actionable insights from data. This includes generating scores, forecasts, suggesting "next best actions," or identifying patterns. These derived insights are then integrated into existing processes to provide recommendations. This type of AI represents a demanding layer because it transforms raw material into meaningful signals. For Signal Engines, rigorous data quality is paramount. Data must be clean, structured, and machine-readable; a forecast built on incomplete or contradictory historical data will be inaccurate and ultimately worthless. Projects in this category absolutely require a solid data foundation to be successful.

### In-Process Assistance

The "in-process assistance" category of AI focuses not on prediction but on providing information and overview to users during their work. For instance, an employee preparing a quote or handling a customer interaction might need specific information scattered across various systems, documents, and inboxes. AI in this context finds, summarizes, and presents this information to aid decision-making. Here, the data's unstructured and scattered nature is precisely why AI is valuable. The disorder is not an impediment but the very problem AI is designed to solve. Harmonizing this data beforehand would eliminate the need for such AI assistance.

### Why the Confusion Is Expensive

The indiscriminate application of the "data first" mantra leads to costly delays. Consider a company wanting to equip its sales team with AI to quickly access historical terms, technical specifications, and conversation logs during the quoting process. This is a clear case for "in-process assistance," a use case that is immediately feasible. However, due to the reflexive application of the "data first" principle, such initiatives often devolve into multi-year data harmonization projects. This postpones value delivery for years, even though the specific AI application would not have required such a data foundation. While clean data always seems like a logical prerequisite, for in-process assistance, such a preliminary data migration is a costly detour.

### The Real Investment Question for CIOs

This distinction does not negate the importance of data quality entirely, but it calls for a more nuanced approach to prioritizing investments. For "in-process assistance" use cases, the immediate goal should be to assist users in their work, regardless of data fragmentation. The scattered, unstructured nature of the data is the core problem the application addresses. Conversely, for "Signal Engine" applications, the true bottleneck and most demanding investment is not the AI model itself, but the underlying data layer that transforms heterogeneous sources into reliable signals. CIOs should not ask "Is our data ready?" but rather, "Which AI world does this use case belong to?" This question dictates whether data quality is a primary hurdle or irrelevant.