← The Blog

Replacing Databricks Compute with Falcon: Frequently Asked Questions

Straight answers to the questions data engineering and platform teams ask before moving production compute off the JVM.

Most teams evaluating Haevek's Falcon Compute Platform arrive with the same set of questions: what breaks, what stays, what access we need, and how long it takes to find out. Below are the answers, drawn from customer benchmarks and production deployments.

For the full argument behind these answers, see Slash Your Databricks Compute Bill by over 80 Percent Without Migrating Anything. For the operational sequence, see The Databricks Offload Playbook.

Platform and Architecture

What is the Falcon Compute Platform?

Falcon is a distributed compute engine built in Rust as a complete ground-up redesign of Apache Spark. Every data processing step compiles to machine instructions that execute directly on the CPU, with optimal offloading to specialized chips (GPUs, TPUs, and similar) should workloads require it. It runs Kubernetes-native: each pipeline executes as an atomic K8s job, so compute activates when the job starts and releases when it finishes.

Does replacing Databricks compute require migrating data or changing Unity Catalog configurations?

No. Falcon is a pure compute engine that reads from and writes to your existing Delta Lake or Iceberg tables in your current object store. Unity Catalog remains the governance layer with no re-registration, no access policy changes, and no table contract modifications. Your lakehouse storage layer is completely untouched.

Which Kubernetes environments does Falcon support?

Falcon deploys on any flavor of k8s including: EKS (AWS), AKS (Azure), GKE (Google Cloud), and OpenShift. It installs into your existing Kubernetes environment inside your own cloud account or on-premise environment — no new accounts or storage layers are required.

How does Falcon handle AI inference workloads alongside batch ETL?

Falcon supports direct bindings to existing libraries across languages, so teams can call Python, C/C++, or other dependencies natively within a pipeline without a separate service layer. As one example of what this enables, ONNX (Open Neural Network Exchange, an open standard format for packaging and sharing machine learning models) models can be embedded directly in the pipeline, eliminating the separate model-serving infrastructure many teams run alongside Databricks for batch inference. This reduces both operational complexity and the number of compute footprints to manage.

Cost and the JVM

Why don't standard Databricks optimizations like Photon or auto-termination reduce the bill enough?

Photon can reduce job runtime for certain analytical queries, but Databricks charges substantially more DBUs for Photon-enabled compute, making the cost impact largely net neutral. It also doesn't support every operation a pipeline might use, so many jobs only partially benefit from it even when it is enabled. Auto-termination only affects idle interactive clusters, which represent 20–30% of a typical enterprise bill. The 70–80% sitting in "always-on" production compute is structurally unaffected by either optimization.

What is the JVM overprovisioning multiplier and why does it matter for Databricks cost?

The JVM interprets Java bytecode, meaning every operation passes through an abstraction layer rather than compiling directly to machine instructions. When a job processes terabytes of data, the inefficiency of the abstraction layer has measurable impacts. For example, it constantly serializes records into JVM objects, ships those objects between workers across the network, and deserializes them on the other side. Another impact is that the JVM's garbage collector periodically halts all threads to reclaim memory — a behavior known as stop-the-world GC. Since these pauses are unpredictable, teams typically provision significantly more compute than a workload theoretically requires just to reliably hit SLA windows. The result is a structural cost floor that cluster tuning and right-sizing can reduce at the margins but can't eliminate. Switching to a compiled native runtime removes that floor entirely.

Does switching to a managed alternative like EMR or Dataproc solve this?

No. Whether you run a managed service (Databricks, Amazon EMR, Azure HDInsight, Google Dataproc) or self-managed Spark, the underlying cloud compute bill stays the same. What varies is what you pay on top: managed services add a licensing layer, while self-managed Spark trades that licensing cost for the engineering hours needed to maintain and tune the cluster yourself. None of those choices reduce the compute footprint itself, because the JVM inefficiency runs underneath all of them.

How much can a team realistically expect to save?

It depends on how much of your bill is production compute and how heavily you're overprovisioned. For a representative mid-market account with a $400K annual Databricks bill where 75% is production compute, that is $300K running on the JVM that can be migrated. Applying the 85% TCO reduction observed in the commercial streaming benchmark — a figure that already accounts for Falcon's license cost — yields roughly $255K in annualized savings, with the remaining $100K of interactive work staying on Databricks untouched. Teams running 5–10x excess capacity see the largest absolute savings, because the efficiency gain applies to the biggest base.

Benchmark · Commercial data analytics company

432 cores on Databricks. 32 cores on Falcon. A 93% reduction in compute infrastructure.

A production real-time streaming ETL workload processing multi-terabytes per day. Same transforms, same data. Annual spend dropped from $464K to $66K for this job, an 85% reduction in total cost of ownership.

Evaluation and Scope

How long does a Falcon benchmark take, and what access does it require?

The benchmark window runs 2–4 weeks for the side-by-side measurement phase. Falcon needs narrowly scoped read access to specific source object store paths and write access to a designated output path. No access to your Databricks workspace or Unity Catalog governance layer is required.

Which workloads should stay on Databricks?

Interactive SQL exploration, notebook-based ML experimentation, Delta Live Tables pipelines tightly coupled to Databricks orchestration, and anything where analyst iteration speed matters more than infrastructure efficiency. The goal is not to leave Databricks. It is to shift the 70–80% of spend running on production compute off the JVM runtime and keep Databricks where it earns its place.

How do I identify which of my jobs to move first?

Query system.billing.usage over a 30-day window and sort descending by total DBUs. Flag anything running more than four hours per execution, or anything with a run count near 720 for the month — that's an hourly job, which means it is always on. Three categories are typically worth targeting: long-running batch ETL, 24/7 streaming ingestion, and high-throughput analytics or ML inference. The full query is in the main article.

How is the cutover sequenced so we don't take an outage?

Shadow mode first: Falcon runs alongside the existing pipeline on the same input with no production traffic routed to it, so you can compare outputs and benchmark cost before anything shifts. Once output parity and performance targets are confirmed, you route production traffic to Falcon and keep the rollback path live until you are satisfied that it can be shut down. For API-driven workloads, an optional canary phase shifts 5–10% of traffic incrementally until all traffic is migrated over to Falcon. The full sequence is in The Databricks Offload Playbook.

Question not answered here?

Email info@haevek.com with the subject line "Test Flight."

Share LinkedIn X Email

Take a Test Flight

Same workload. Less infrastructure. Less time.

Take a Test Flight →