Nobody Would Be Stupid Enough to Rebuild Apache Spark (So We Did)

Written by Haevek | Aug 12, 2026, 8:37:20 PM

The founding story of the Falcon engine: two failed versions, a language decision made over a winter holiday, and the architectural conflict that took three rewrites to resolve.

There's an article arguing that no startup would be stupid enough to rebuild Apache Spark from scratch. The reasoning is sound: a decade of development history, thousands of contributors, and a scope that would exhaust any startup's runway well before producing anything useful.

We found that article while we already had a working version.

Getting there took longer than we expected and required solving problems we hadn't anticipated. This is the story of what those were.

The founding problem surfaced twice, from different directions. The first was a government program running a 15-year-old distributed compute system written in C++ — fast enough to meet processing requirements, but so complex that it needed several dozen engineers with PhDs and two decades of C++ experience just to keep it running. Apache Spark seemed like the obvious path to a leaner architecture. The engineers on the program had already tried it. It wasn't fast enough. The second was a rare event prediction program where dev environments alone were running $100,000 to $200,000 a month in compute before anything reached production. The customer was direct: you can't keep throwing infrastructure at this problem. Both situations pointed to the same conclusion: Spark needed to be faster, and the only way to get there was to rebuild the execution engine itself.

That conclusion was easy to reach and uncomfortable to sit with. Rebuilding it in C++ was the obvious move technically, but C++ puts the entire burden of memory safety on the developer. The pattern is well documented across production systems: memory mismanagement accounts for the majority of exploitable vulnerabilities in software built in C++, and the maintenance cost follows. A team that had spent years cleaning up poorly written C++ code, tracing memory leaks with unreliable tooling and rewriting implementations that worked on paper but couldn't be sustained in production, had no interest in building that problem into a new foundation. We spent a winter holiday benchmarking several major backend systems languages available: C, C++, Go, Python, Rust, and Java. Rust matched or exceeded C++-level performance across every test that mattered. It enforces memory safety at compile time through strict ownership and borrowing rules, with no garbage collector, which means no GC pauses and no runtime overhead. Go was close on performance but came with a garbage collector, which rules it out for production compute workloads requiring consistent throughput. The choice was straightforward once the data was in front of us.

NSA and CISA reached the same conclusion independently. Their joint guidance published in June 2025 recommends memory-safe languages as a primary strategy for reducing software vulnerabilities. Google Project Zero found that 75% of CVEs exploited in the real world were memory safety vulnerabilities. Android reduced its share from 76% to 24% by switching new development to Rust. For Haevek's Department of Defense, Intelligence Community, and NATO customers, building in Rust meets both a performance requirement and a compliance mandate their own agencies are pushing the industry toward.

The deployment architecture came from a separate insight: a piece on Medium arguing that the JVM was dead and the new virtual machine was Linux in a container on Kubernetes. Years of deploying large-scale software into air-gapped government environments, getting systems to bootstrap without remote access and work through authority-to-operate processes for classified deployments at an enterprise AI platform company, had made clear what Kubernetes could do. If compiled Rust code in a container carried essentially no overhead, the same compute engine could run in a hyperscaler, in a customer data center, or on an edge node with no internet connection. We validated that overhead assumption the same winter holiday. It was unmeasurable at the scales we were operating at.

Then we started building the engine, and the first two versions failed.

The problem that drove both failures, and that took a third attempt to resolve, is a fundamental architectural conflict. High-performance compiled languages require static typing: you define at compile time how the whole system is going to work, and that's where the performance comes from. But a general-purpose distributed compute engine has to handle any workload, with different data types, different pipeline shapes, different compute profiles, and different source and destination formats. Those two requirements pull directly against each other, and working out how to satisfy both simultaneously while keeping the system usable for developers who would build their jobs on top of it is what three full rewrites produced. The techniques developed to resolve that conflict are now the foundation of Haevek's patent portfolio, and they represent years of work that any team attempting this independently would have to produce on their own.

The Bauplan team made the same assessment in their 2023 paper on building a serverless Data Lakehouse, describing a from-scratch rebuild as "a challenge unfit for a startup" and choosing to reuse existing components instead. For most infrastructure problems, that's the right approach. For what we were trying to do, Spark's performance ceiling isn't addressable through configuration or workload-level optimization. It's architectural. By the time we found the article saying nobody would attempt this, we'd already had a working version for some time.

A dozen person-years later

Falcon processes the same distributed workloads as Spark at up to 25 times faster performance, and deploys anywhere Kubernetes runs.

The engine already covers approximately 70 to 80 percent of the functionality data engineering teams use day to day. The first customers to put Falcon into production replaced their Databricks workloads in eight days.

In one case, running against compiled code written for a government program, Falcon measured up to six times faster performance on certain operations. Two factors likely explain this: Falcon's resource allocation model uses Kubernetes to go well beyond the standard one-thread-per-CPU assumption and aligns compute footprint to whether a workload is network-bound, disk-bound, or compute-bound. Developers working in compiled languages tend to optimize the inner loop while treating serialization, deserialization, queuing, and cache management as details the language will handle. Those aren't small inefficiencies at scale, and significant engineering effort has gone into each of them precisely because that's where performance gets lost when nobody's paying attention.

If you want to see what the numbers look like on your workload rather than ours, that's what the Falcon test flight is for. The Falcon test flight takes a real workload from your environment, runs it side by side against your existing implementation on equivalent infrastructure, and delivers the result before any commercial commitment. The claims in this article were earned the same way.

Get started

Contact Haevek to learn more or sign up for a test flight.

Keep reading

Slash Your Databricks Compute Bill by over 80 Percent Without Migrating Anything
Why 70–80% of your bill sits in production compute, and how to find your offload candidates.

Spark Is Free. Running It Isn't.
Why the free license and the infrastructure bill are two different conversations.