Talk

PyData Amsterdam 2026: Adyen and LakeSail on Modernizing Spark with Sail

Why Adyen went looking for a new engine, and what happened when Sail went into production.

Santosh Pingale, Principal Engineer at Adyen, and Shehab Amin, co-founder and CEO of LakeSail, at PyData Amsterdam 2026.

Adyen’s data platform handles more than ten thousand Spark jobs a day, across fifty-plus teams and more than a million lines of PySpark, on-prem and at petabyte scale. The company processes more than a trillion dollars for its merchants. Sail’s output had to match Spark’s exactly. Each job ran alongside its Spark version, with schema and data equality checked in Spark. Missing features went upstream, including Hive Metastore support. Sail now runs in production.

They explored DuckDB and Polars before settling on Sail. Both would have meant a second API across the platform, two sets of utilities and integrations to maintain, rewrites and validation for every job that moved, ecosystem gaps starting with HDFS, and another migration the first time a job outgrew one node. Tuning Spark meant three-hour feedback loops. Accelerators stay bounded by the engine underneath them. Sail keeps the Spark API and replaces the JVM, so the platform built on that API stayed as it was.

Results

The jobs Adyen ran on Sail finished in a fraction of Spark’s time. A CI test suite that took 1,335 seconds on JVM Spark ran in 78.44 seconds on Sail, about 17x faster. The result was fast enough that the engineer who ran it thought it was a bug. On an Airflow job mentioned in the presentation, Spark’s median runtime was about 57 minutes and Sail’s was about 8.5 minutes, about 6.7x lower, on fewer cores and less memory. Be sure to watch the talk for more details.

In the Talk

  • 4:26 The fixed cost of a small Spark job, and why the alternatives fell short.
  • 8:58 What Sail is.
  • 16:47 The rollout method.
  • 19:38 "At first, I thought it was a bug."
  • 20:50 Production results.
  • 21:22 Ecosystem compatibility, not just API compatibility.
  • 22:11 Q&A on Spark 3.5 and 4 support, moving off RDDs, and migrating from managed platforms.

Try Sail Today

Sail is open source at github.com/lakehq/sail. Run pip install pysail, start a server, and point an existing PySpark job at it. The Getting Started guide covers the setup. If you would rather not operate it yourself, the LakeSail Platform is managed Sail in your own cloud, with a free trial.

Migrate with Guaranteed Savings

LakeSail will prove one of your recurring Spark workloads on Sail at no service fee, then migrate it against a scope agreed in writing. If the result is not correct, faster, and cheaper, you owe nothing for the engineering. Please reach out and learn more about our Spark Migration Guarantee.

Questions about your Spark setup?

Book 30 minutes with a LakeSail engineer. Real answers, no pitch.