Skip to content

OneLake in Microsoft Fabric

One Data Layer Instead of Multiple Copies and Systems

Distributed Data Architecture

In traditional data architectures, each layer – whether a data warehouse, data lake, or BI system – often operates on its own copy of the data. This increases complexity, costs, and makes management more difficult. Microsoft Fabric introduces a different approach. Instead of multiple distributed data stores, it relies on a single, shared layer- OneLake.

Architektura OneLake w Microsoft Fabric oparta na Azure Data Lake Storage Gen2 i Delta Lake

What is OneLake

OneLake is a shared data layer within Microsoft Fabric that is used by all platform components – from data engineering to reporting in Power BI.

It serves as a central data repository for the entire organization, eliminating the need to maintain multiple copies of the same data across different systems. All processes in Microsoft Fabric – data engineering, analytics, and data science – use the same data layer.

This approach is based on the concept of a single source of truth. In practice, this means reducing data silos, simplifying management, and ensuring consistency for teams working on the same datasets.

OneLake is not a separate data storage system, but a logical layer built on Azure Data Lake Storage Gen2. Data is stored in open formats such as Delta Lake, enabling it to be used by different analytical engines without the need for migration or transformation.

How OneLake Works?

Data format

Azure Data Lake Storage Gen2 jako warstwa przechowywania danych w OneLake

OneLake uses Azure Data Lake Storage Gen2 as its underlying storage layer, enabling data to be stored in a single place – from structured to unstructured data.

OneLake stores, among others::

Structured Data – Delta tables, Parquet files, CSV

Semi-structured-data – JSON, XML

Unstructured data – text files, images, logs, and documents

Delta Lake jako format danych zapewniający transakcyjność i wersjonowanie w OneLake

A key element of OneLake is the Delta Lake format.

Delta Lake is an open table format built on Parquet, extended with a transactional layer (ACID) and data versioning. This enables safe data processing during concurrent operations and access to historical versions of the data.

Key features of Delta Lake:

ACID transactions – ensure data consistency during concurrent operations

Time travel – access to previous versions of data

Schema enforcement – control over data consistency with the defined schema

optymalizacje wydajności – improved data organization and read efficiency

Shortcuts

Shortcuts in OneLake allow you to use data located outside Microsoft Fabric as if it were part of the same data space.

Instead of copying data between systems, OneLake lets you create logical references to external sources. This way, the data remains in its original location while still being accessible to all Fabric components.

This approach simplifies the data architecture and reduces the number of synchronization operations. Changes made at the source are immediately reflected in OneLake, without the need for additional update processes.

Security

Microsoft Entra ID jako system zarządzania dostępem i tożsamością w OneLake

Access to data in OneLake is managed through Microsoft Entra ID, which handles user authentication and permission control.

Access control is applied at multiple levels:

Workspace

Item

Folder and data

This allows precise definition of who can access specific data and perform operations. Different teams can work on the same data layer while having access only to selected elements.

Additionally, the role-based model limits user actions:

Admin – full control over the workspace

Member – edit and manage resources

Contributor – create content

Viewer – read-only access

OneLake also supports secure communication with data sources via mechanisms like Private Link or Data Gateway, helping to limit public data exposure and maintain control over information flow.

Support for multiple tools and technologies

Thanks to the shared data format in OneLake, the same data can be used by different analytics engines without the need for copying.

In practice, this allows working on a single data layer using various approaches and technologies:

Apache Spark – data processing and data science (PySpark, Scala, Spark SQL)

SQL (Polaris) – analytical queries using T-SQL

Python – data analysis and quick operations without starting a cluster

Direct Lake – direct access from Power BI to data in OneLake

KQL – real-time data and log analysis

Each of these engines operates on the same data, eliminating the need to create separate layers for different use cases and simplifying the overall analytics architecture.

OneLake as the foundation of Microsoft Fabric

OneLake serves as the central data layer for all Microsoft Fabric components. Whether data is processed by a data engineering team, analyzed in Power BI, or used in data science models, all operations occur on the same data layer.

In traditional architectures, individual tools often rely on their own data stores or copies of data. Microsoft Fabric simplifies this approach – every platform component operates directly on the data stored in OneLake.

This approach eliminates the need to replicate data across systems, streamlines data flow, and enables multiple teams to work in parallel on the same datasets.

OneLake jako centralna warstwa danych dla komponentów Microsoft Fabric

Data Source Integration

OneLake enables direct work with data from various systems without the need to copy or duplicate it in multiple locations. The platform integrates cloud sources, on-premises systems, and SaaS services, creating a single, unified data layer accessible across the entire organization.

DThis allows teams to use the same data regardless of the tool or context – whether for analytics, reporting, or AI models.

Instead of building complex ETL processes and maintaining multiple copies of the same data, OneLake allows direct work on the sources through mechanisms such as shortcuts. Data stays in its original location while remaining accessible within a single architecture, significantly simplifying management and reducing maintenance costs.

OneLake jako centralna warstwa danych integrująca różne źródła w Microsoft Fabric

Single Data Layer, Real Business Advantage

OneLake transforms the way organizations work with data – rather than managing multiple copies and distributed systems, all teams operate on a single, unified data layer.

This means reduced complexity, lower maintenance costs, and faster access to up-to-date information. Data is no longer locked in silos- it becomes available wherever it is needed, in real time.

As a result, companies can make decisions more quickly, better leverage analytics, and build data-driven solutions without the limitations typical of traditional architectures.

Contact Us

Do you want to organize your data and reduce duplication across your organization?
Schedule a short consultation to see how OneLake can simplify your data architecture.

Antdata - calendar person