---
title: "Lakehouse: The New AI Infrastructure"
author: "Daniel Gorld"
author_role: "Consulting Director, cbs CX — The cbs Group Salesforce Consultancy"
author_url: "https://cx-waves.com/about"
publisher: "CX-Waves"
canonical_url: "https://cx-waves.com/nodes/lakehouse-ai-infrastructure"
date_published: 2026-06-21
date_modified: 2026-09-20
language: en
---

# Lakehouse: The New AI Infrastructure

Source: Daniel Gorld, CX-Waves — https://cx-waves.com/nodes/lakehouse-ai-infrastructure (published 2026-06-21, updated 2026-09-20)

## Why is the data lakehouse becoming a critical piece of enterprise AI infrastructure?

*   **Enables Dual Consumption:** A lakehouse allows a single dataset to serve two distinct consumers: an analytical "signal layer" for statistical predictions and an "AI agent" for process-oriented contextual actions, optimizing data utility.
*   **Addresses ERP Data Access Restrictions:** As major ERP vendors tighten control over direct data access from their core systems, the lakehouse provides a strategic escape route, allowing companies to maintain control over their data for AI initiatives.
*   **Prioritizes Data Governance:** Recent vendor partnerships emphasize cross-platform authentication, identity mapping, and metadata-aware access controls, signaling that governance, not just availability, is the new bottleneck for productive AI agents.
*   **Shifts Control to the Enterprise:** By moving relevant data into a company-controlled analytical home, organizations retain authority over their strategic data assets, rather than depending solely on vendor-controlled integration paths.

## Two Very Different Consumers of the Same Data Set

A lakehouse, which functions as an analytical data home integrating structured and unstructured data, serves two distinct categories of consumers from a single underlying dataset. This architectural pattern accommodates varying data needs within an enterprise.

The first type of consumer is the "signal layer," responsible for computing machine learning-based predictions. This includes identifying customer churn, detecting equipment wear patterns, or assessing the probability of sales deal closures. The effectiveness of this layer is directly tied to the quality and consistency of historical data.

The second consumer is the "agent in the process," which requires searchable and reliable context rather than perfect statistical foundations. This agent utilizes information such as order history, technical specifications, service cases, or contractual status. Through zero-copy data integration, like that offered by CRM platforms, this contextual data is made available to the agent without duplication or separate maintenance.

The critical insight here is that the same dataset supports two different logics of use. Conflating these demands—insisting on the same high data quality for every agent as for a predictive model—can lead to misdirected investments. Maintaining a clear separation between these data use cases is a process-driven decision, not purely a technical one.

## Why Now — the Pull at the ERP Core

The urgency for this architectural approach, particularly for customers using SAP landscapes, stems from a changing dynamic with market-leading ERP systems. These systems are increasingly restricting access to their data. Traditional methods for third-party systems to extract data from the ERP core are being narrowed through both contractual terms and technical implementations.

This trend indicates that ERP vendors prefer external agents to interact with their systems through vendor-controlled pathways, rather than directly accessing the productive core. This is where the lakehouse transforms from a technical design concept into a strategic imperative for businesses.

For companies unwilling to solely rely on ERP vendor roadmaps, the lakehouse offers a solution: migrating relevant data consistently into a dedicated analytical environment. This enables evaluation of the data and its provision as context for agents, preventing direct interaction with the productive core. The control over data shifts from the tightly managed source system to a data home managed by the company itself.

While vendors could also impose restrictions on data offloading, historical evidence suggests that preventing companies from accessing data essential for their operations proves unsustainable in the long term. Thus, the lakehouse serves as a consistent answer to concerns about data dependency: when access to core systems becomes more restrictive, companies can secure their data context in an independently controlled environment.

## The Process Decides, Not the Store

Despite the architectural advancements, a crucial limitation persists, often overlooked amidst the enthusiasm for "AI memory." While a lakehouse functions as a long-term memory for AI, memory alone does not equate to action. A functional AI architecture requires interplay: the CRM provides process context, the lakehouse supplies memory, and the agent performs actions. Value is generated only through this combined effort.

Furthermore, even the most comprehensive dataset does not automatically become useful context. Defining which data an agent can access in specific process situations, establishing relevant permissions, and determining what information is required for particular business events involves significant modeling and process design. These are not merely questions of data storage location. The lakehouse facilitates data availability; however, whether this leads to reliable signals or correct actions is ultimately determined by the processes layered on top of it.

For industrial companies aiming to scale their AI initiatives beyond initial pilots, the core message is clear: the focus should not solely be on new data-lake components. Instead, the infrastructure providing memory and context for AI is actively evolving, and its effective utilization necessitates a process-driven perspective, rather than one centered solely on data storage.