Microsoft OneLake gives every Microsoft Fabric tenant a single, unified data lake built on open formats. One copy of data, many engines: that is the promise, and it is what makes OneLake such a strong foundation for analytics and AI alike. This post looks at how Sail, the open-source distributed compute framework, and LakeSail Cloud, the managed platform for running Sail in production, connect natively to OneLake, and what that unlocks for teams running Spark and Python workloads on Fabric.
Sail is a drop-in replacement for Apache Spark, rebuilt from the ground up in Rust, with Rust-native support for Delta Lake and Apache Iceberg out of the box. That makes it a natural fit for OneLake, where lakehouse tables are stored in exactly these open formats. Point Sail at a Fabric lakehouse and it reads and writes the same tables the rest of Fabric uses, with no data copies, no extra connectors to install, and no changes to existing Spark code.
One Engine for Data and AI
Sail implements the Spark Connect protocol, the gRPC interface that decouples Spark clients from the engine that executes their plans. Existing Spark SQL and DataFrame code targeting Spark 3.5 or 4.x runs as is from any Spark Connect client, whether PySpark, Scala, Go, or Rust. The only thing that moves is the connection string.
from pyspark.sql import SparkSession
spark = SparkSession.builder.remote("sc://sail-host:50051").getOrCreate() Behind that familiar interface sits a very different engine. Sail is written entirely in Rust with no JVM. Data stays in Apache Arrow columnar format end to end; query execution builds on Apache DataFusion with vectorized SIMD kernels, and an async actor runtime coordinates distributed work without locks.
When it comes to Python, Sail treats it as a first-class citizen. In Apache Spark, Python UDFs run in separate worker processes, so every batch of data is serialized, shipped across a process boundary, deserialized for Python, and then sent back the same way. Sail removes that boundary entirely: the Python interpreter is embedded in the engine itself, so UDFs execute in-process and read the same Arrow buffers as the rest of the query, with no copies and no serialization step. Python becomes a native part of the execution plan rather than a detour out of it, which matters for AI workloads, where heavy lifting so often happens in Python.
The same engine runs in two deployment modes. On Kubernetes, it runs as a distributed cluster for production scale. On a single node, whether a laptop or one large VM, it is a great fit for local development, testing, and smaller datasets, and it works out of the box on Windows, macOS, and Linux: with no JVM to install or configure, a pip install is the whole setup. Either way, a Sail server starts in seconds, and its workers are stateless and lightweight, so a cluster can scale out quickly when load arrives and back down to zero when it passes.
The architecture shows up in the numbers. In a derived TPC-H benchmark, Sail completed the full run in 52.81 seconds versus 534.78 seconds for a self-managed Apache Spark deployment on the exact same hardware, with per-query speedups ranging from 176% to 2819%. Sail’s peak memory was 26 GB, compared to 72 GB for Spark, and Sail wrote zero shuffle spill to disk while Spark wrote about 115 GB. In practice that means the same workload can run on an instance about a quarter of the size, and since the roughly 10x speedup also cuts runtime to a tenth, the combined effect is one fortieth of the compute cost, a nearly 98% reduction. Benchmarks are always workload dependent, so the best evaluation is running your own pipelines, but the pattern is consistent: less memory, less disk, less time.
The Workloads This Serves
The integration is aimed at patterns that will look familiar to most data teams on Fabric:
- Scheduled ETL. Nightly PySpark jobs that transform raw orders, events, or clickstream data into curated tables in a lakehouse. The code already exists; the goal is to run it faster and cheaper without a rewrite.
- Interactive exploration. Analysts and engineers querying production-scale tables from a notebook or IDE, where a tight iterate-and-refine loop matters more than anything.
- AI enrichment. Generating embeddings over product reviews for semantic search, classifying support tickets with a language model, or running feature engineering that goes beyond built-in SQL functions. These steps live in Python, and they increasingly sit in the middle of ordinary data pipelines rather than off to the side.
- Agent-driven analytics. AI agents that explore tables, run queries, and validate results on a person’s behalf. Sail ships an MCP server so agents can query lakehouse data conversationally, plus a one-shot CLI: pipe a PySpark script into
sail spark runand a local server starts instantly, executes the script, and shuts down when it finishes. The Sessions capability covered below gives each agent its own scoped, revocable access to real compute.
The common thread: Spark APIs and Python are the skills, OneLake is where the data lives, and the pressure on cost and latency grows as AI steps enter the pipeline. Because Python runs inside the Sail engine rather than beside it, the same engine that handles the SQL handles the model calls, and the UDF stops being the slow part of the job.
Try Sail Open Source with OneLake
Getting started takes a few minutes. Install Sail from PyPI:
pip install "pysail" Sail ships a OneLake catalog provider that talks to OneLake’s open table APIs. It supports two modes: iceberg, which uses the Iceberg REST catalog endpoint, and delta, which uses the Delta Lake endpoint. For the full list of supported authentication methods, see the OneLake catalog guide in the Sail documentation.
The rest fits in one short Python script: set the credentials, register a Fabric lakehouse as a catalog, start a Sail server in-process, and connect to it. The catalog url follows the workspace/item-name.item-type convention, so a lakehouse named sales in a workspace named analytics becomes analytics/sales.Lakehouse. In a schema-enabled lakehouse, tables sit under the dbo schema by default:
import os
from pysail.spark import SparkConnectServer
from pyspark.sql import SparkSession
# Authenticate with a Microsoft Entra service principal
os.environ["AZURE_TENANT_ID"] = "<tenant-id>"
os.environ["AZURE_CLIENT_ID"] = "<client-id>"
os.environ["AZURE_CLIENT_SECRET"] = "<client-secret>"
# Register the lakehouse as a catalog named "fabric"
os.environ["SAIL_CATALOG__LIST"] = (
'[{type="onelake", name="fabric", url="workspace/lakehouse.Lakehouse", api="iceberg"}]'
)
# Start a Sail server in-process and connect to it
server = SparkConnectServer()
server.start()
_, port = server.listening_address
spark = SparkSession.builder.remote(f"sc://localhost:{port}").getOrCreate()
spark.sql("SHOW TABLES IN fabric.dbo").show()
orders = spark.table("fabric.dbo.orders")
daily_revenue = (
orders
.groupBy("order_date", "region")
.sum("amount")
.withColumnRenamed("sum(amount)", "revenue")
)
daily_revenue.write.format("iceberg").saveAsTable("fabric.dbo.daily_revenue") The write lands in OneLake as a regular lakehouse table. It is immediately visible in the Fabric portal and immediately usable everywhere else in Fabric: Power BI through Direct Lake, Fabric notebooks, the SQL analytics endpoint, and any other engine reading the same copy of the data. Switching api="iceberg" to api="delta" connects the same lakehouse through the Delta Lake catalog endpoint instead, since Sail supports both formats natively.
AI Enrichment with Arrow Python UDFs
Sail supports the full Spark UDF surface: scalar UDFs, UDAFs, UDWFs, and UDTFs, in plain Python, pandas, and Arrow variants, including mapInArrow and applyInArrow. Because Python executes in-process, Arrow-native UDFs receive Arrow arrays directly from the engine with zero copies, which makes them a natural bridge between ETL and AI.
Here is an Arrow Python UDF that generates sentence embeddings over product review text stored in OneLake, using an open-source model. Each invocation receives a whole Arrow batch of strings, so the model can encode them together instead of row by row:
import pyarrow as pa
from pyspark.sql.functions import arrow_udf
from pyspark.sql.types import ArrayType, FloatType
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
@arrow_udf(returnType=ArrayType(FloatType()))
def embed(texts: pa.Array) -> pa.Array:
vectors = model.encode(texts.to_pylist(), batch_size=256)
return pa.array([v.tolist() for v in vectors])
reviews = spark.table("fabric.dbo.product_reviews")
embedded = reviews.withColumn("embedding", embed("review_text"))
embedded.write.format("iceberg").saveAsTable("fabric.dbo.product_reviews_embedded") The job reads from OneLake, runs Python at engine speed, and writes results back to OneLake, where the embeddings can power semantic search and the categories can land straight in a Power BI report.
Secure, On-Demand Access with LakeSail Cloud Sessions
Sail is open source and can be self-hosted anywhere, including on Kubernetes. LakeSail Cloud runs it in production as a managed service inside your own cloud account, and adds a capability called Sessions that pairs especially well with OneLake.
A production cluster is typically private: no public address, no inbound connections from the internet. A Session provides a way through that boundary without loosening it. From LakeSail Cloud, you mint an access-scoped, time-boxed token, signed with a key unique to your cluster. A hardened, stateless proxy validates the token and forwards traffic only to the one workload the token is scoped to. The cluster never gets a public address, and the data never leaves your cloud account. Compute comes up on demand when a Session opens and is reclaimed when it goes idle, so nothing sits warm between tasks.
Connecting a Session to OneLake takes three steps:
- Create a catalog. In LakeSail Cloud, add a catalog of type Microsoft OneLake, provide the
workspace/lakehouse.LakehouseURL, choose the Iceberg REST or Delta Lake API, and supply credentials. - Create a session. Bind it to a cluster and set the OneLake catalog as its default, then generate a token from the session detail page.
- Connect from anywhere. Point any Spark Connect client at the endpoint:
from pyspark.sql import SparkSession
spark = SparkSession.builder.remote(
"sc://<grpc-endpoint>/;token=<jwt>"
).getOrCreate()
spark.sql("SHOW TABLES IN fabric.dbo").show() That lists the Fabric lakehouse tables through the session’s private cluster. Because the OneLake catalog is the session default, unqualified names also resolve against the lakehouse, so plain SHOW TABLES IN dbo works just as well. A local notebook can explore production-scale OneLake tables interactively, with no sample copied down and no local engine that behaves differently from the real one. A scheduled pipeline or CI step can mint a short-lived token, run a UDF-heavy job like the embedding example above, and finish without any long-lived credential baked into the runner. When the token expires, the gate stops honoring it.
Sessions were also built with agents in mind, a theme that will resonate with anyone following the growing role of AI agents in data work. A token a client can carry is a token an agent can carry. Handing a Session to a coding agent gives it real compute against real OneLake data under bounded trust: one workload, one owner, a fixed lifetime, and no way to reach anything else or extend its own access. Ten agents can each hold their own scoped Session and work in parallel, each one individually attributable and individually revocable.
Better Together
OneLake’s open-format foundation means the same copy of data can be served by many engines, each chosen for what it does best. Sail adds a Rust-native, Python-first option to that ecosystem: drop-in Spark compatibility for the code teams already have, in-process Python for the AI work they are adding, and a performance and cost profile that leaves more budget for building. Everything it writes back to OneLake is instantly discoverable, governed, and consumable across Fabric, from Direct Lake reports to notebooks to the SQL analytics endpoint. One copy of data, many engines, and now one more way to move fast on it.
Get Started
- Read the OneLake catalog guide in the Sail documentation
- Explore Sail on GitHub and the getting started guide
- Follow the Sessions guide or read the Sessions announcement
- Sign up for LakeSail Cloud to run Sail in production in your own cloud account
- Learn more about Microsoft OneLake
Apache®, Apache Spark™, Apache Arrow™, Apache DataFusion™, and Apache Iceberg™ are trademarks (or registered trademarks) of The Apache Software Foundation. Delta Lake™ is a trademark of LF Projects, LLC. LakeSail, Inc. is not affiliated with, endorsed by, or sponsored by The Apache Software Foundation, the Delta Lake project, or LF Projects, LLC.