Concept

Lakehouse: The New AI Infrastructure

The data lakehouse architecture is becoming essential AI infrastructure, enabling a single dataset to serve two distinct consumers: the analytical 'signal layer' for predictions and the process-oriented 'AI agent' for contextual action. This shift is crucial as ERP vendors tighten direct data access, making the company-controlled lakehouse a strategic asset for scalable AI.

Why is the data lakehouse becoming a critical piece of enterprise AI infrastructure?

  • Enables Dual Consumption — A lakehouse allows a single dataset to serve two distinct consumers: an analytical "signal layer" for statistical predictions and an "AI agent" for process-oriented contextual actions, optimizing data utility.
  • Addresses ERP Data Access Restrictions — As major ERP vendors tighten control over direct data access from their core systems, the lakehouse provides a strategic escape route, allowing companies to maintain control over their data for AI initiatives.
  • Prioritizes Data Governance — Recent vendor partnerships emphasize cross-platform authentication, identity mapping, and metadata-aware access controls, signaling that governance, not just availability, is the new bottleneck for productive AI agents.
  • Shifts Control to the Enterprise — By moving relevant data into a company-controlled analytical home, organizations retain authority over their strategic data assets, rather than depending solely on vendor-controlled integration paths.
  • Lakehouse
  • AI Infrastructure
  • Data Governance
  • AI Agents
  • Zero Copy
  • SAP
  • Databricks
  • Salesforce

Auf LinkedIn lesen →

Markdown-Version (zitierfähig)

This week, Salesforce and Databricks expanded their partnership. At first glance, it reads as one announcement among many: two platforms moving closer together, the common denominator being Zero Copy — data made usable across system boundaries without duplicating it. That capability already existed in essence. What is remarkable is less the partnership itself than where the new stage of expansion places its emphasis.

Because the increment does not sit with data access, but with governance: cross-platform authentication, planned identity mapping, metadata-aware access controls. Put differently — the data has long been accessible. What is now being invested in is the question of who may access it, under which rules. That two major vendors devote an entire stage of expansion to this is a signal: the bottleneck for productive AI agents no longer lies in the availability of data, but in the context within which an agent is permitted to act. Otherwise, the investment would not go there.

For industrial companies running an SAP landscape, this is more than a footnote. It makes visible an architecture that will become the quiet foundation of many AI initiatives over the coming years.

Two Very Different Consumers of the Same Data Set

A lakehouse — an analytical data home that brings together structured and unstructured data — serves two entirely different consumers from a single data set.

The first consumer is the signal layer. This is where ML-based predictions are computed: which customer is heading toward churn, which asset shows wear patterns, which quote has the highest probability of closing. This world lives on data quality. A signal is only as good as the history it rests on — here, consistent data maintenance pays off directly.

The second consumer is the agent in the process. It does not need a perfect statistical foundation, but searchable, reliable context: order history, technical specifications, service cases, contractual status. Through the CRM platform’s data integration — precisely this Zero Copy — exactly these contents are made available to the agent without copying them once more and maintaining them elsewhere.

This is the genuinely interesting point: one data set, two logics of use. Anyone who conflates the two — demanding the same data quality for every agent as for a prediction model — invests in the wrong place. Keeping that separation clean is a process decision, not a technical one.

Why Now — the Pull at the ERP Core

There is a second reason this architectural picture is gaining urgency, specifically for SAP customers. A market-leading ERP system is tightening the rules for access to its data. Established extraction paths, through which third-party systems have drawn data from the core for years, are being narrowed both contractually and technically. The direction is recognizable, even if the details are not yet conclusively settled: the vendor does not want external agents working directly on the productive core. Access is meant to run through the vendor’s own, vendor-controlled paths.

This is exactly where the lakehouse turns from an architectural topic into a strategic option. Anyone unwilling to wait for the ERP vendor’s roadmap has a way out: move the relevant data consistently into a dedicated analytical home, evaluate it there, and make it available as context for agents — rather than letting the agent dock onto the productive core. The center of gravity shifts from the controlled source system to a data home over which the company itself retains authority.

This deserves a sober view. A vendor can treat the offloading of data restrictively too, and it would be naive to ignore that. The attempt to prevent extraction entirely, however, has repeatedly failed in practice — data a company needs for its own operations cannot be walled in over the long run. The lakehouse is therefore not a contradiction to the concern about data dependency, but its consistent answer: when access at the core grows tighter, you secure the context at a place you control yourself.

The Process Decides, Not the Store

For all the architecture, one limitation remains — one readily passed over in the current enthusiasm for “the memory of AI.” A lakehouse is a long-term memory, but memory alone does not act. In a workable architecture, the CRM supplies the process context, the lakehouse the memory, the agent the action. Only the interplay produces value.

And even the most complete data set does not become useful context on its own. Which data an agent may see in which process situation, which permission applies, which business event requires which information — that is modeling and process work, not a question of storage location. The lakehouse makes the data available. Whether a reliable signal or a correct action results from it is decided by the process layered above it.

For industrial companies looking to move their AI initiatives beyond pilot status, that is the real message of this week. Not: “there is a new data-lake building block.” Rather: the infrastructure everyone is searching for memory and context in is taking shape right now — and whoever wants to use it should think of it from the process, not from the store.

Daniel Gorld

Consulting Director, cbs CX (The cbs Group Salesforce Consultancy)

Daniel Gorld is a B2B CX and process consultant and Consulting Director at cbs CX (The cbs Group Salesforce Consultancy) with over 20 years of experience in industrial B2B.

Mehr über DanielLinkedIn →

Related

Published · Updated