
In this episode, I sit down with Shehab Amin, co-founder and CEO of LakeSail, to talk about what happens when you decide to rebuild one of data engineering’s most important technologies, Apache Spark, in Rust … sorta.We dig into Shehab’s journey from writing C++ as a kid to studying computer science at Berkeley, becoming a founder, and eventually building Sail. We talk about why Spark has remained so sticky, what Rust, Apache Arrow, and DataFusion are changing about data infrastructure, and why rebuilding Spark compatibility turned out to be much harder—and more interesting—than expected.We also get into the increasingly complicated modern data stack: Delta Lake vs. Iceberg, catalogs, streaming vs. batch, agentic coding, and what happens when AI agents start interacting directly with data and compute.Most importantly, we explore LakeSail’s bigger idea: what if the data pipelines companies already have could become the foundation for their AI pipelines?A wide-ranging conversation about Spark, Rust, AI, and where data engineering goes from here.Thanks for reading Data Engineering Central! This post is public so feel free to share it. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit dataengineeringcentral.substack.com/subscribe
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

The Future of Data Engineering in the Age of AI | Erfan Hesami

Big Tech Laid Him Off. He Realized He Didn’t Need a Job. — Asian Dad Energy

Agentic Data Engineering Is Here — But Can It Close the Loop?

What Happens When a Software Engineer Builds a Company Alone?
Free AI-powered recaps of Data Engineering Central Podcast and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.