<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" version="2.0">
  <channel>
    <title>Haevek Inc. blog</title>
    <link>https://blog.haevek.com/posts</link>
    <description>Notes and updates from the team behind Falcon.</description>
    <language>en</language>
    <pubDate>Mon, 21 Sep 2026 15:18:33 GMT</pubDate>
    <dc:date>2026-09-21T15:18:33Z</dc:date>
    <dc:language>en</dc:language>
    <item>
      <title>Why Haevek Doesn't Use Consumption Pricing</title>
      <link>https://blog.haevek.com/posts/no-consumption-pricing</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Charging on a meter is the wrong incentive.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Charging on a meter is the wrong incentive.&lt;/p&gt;  
&lt;p&gt;Most data infrastructure vendors charge by consumption. The vendor tracks each query, each compute hour, each gigabyte. At the end of the month they issue a bill for what was used. On the surface it sounds fair: you only pay for what you use.&lt;/p&gt; 
&lt;p&gt;Falcon, Haevek's distributed compute engine, doesn't work that way. Our customers pay a predictable fixed annual fee. We built the business that way on purpose, and it is worth explaining why.&lt;/p&gt; 
&lt;p&gt;We charge a fixed annual fee based on a tiered schedule of the maximum number of concurrent cores. A customer picks the capacity they need and pays the same price for the year. They could be running highly tuned workloads or less efficient pipelines they just threw together. It does not affect what they pay us. This is good!&lt;/p&gt; 
&lt;h2&gt;How consumption pricing became the default&lt;/h2&gt; 
&lt;p&gt;Consumption wasn't always how the industry sold software. For decades, companies bought software on subscriptions, licenses, and annual fees — fixed costs. Cloud infrastructure, starting with AWS in the early 2000s, changed that. Once AWS made pay-as-you-go the way to buy compute, similar consumption pricing models appeared for software built on top of cloud infrastructure. Data platforms adopted it especially fast, since usage is easy to meter. By the early 2020s, consumption pricing wasn't just common, it was assumed. New products launched with consumption pricing, and even companies that had built their business on subscriptions came under pressure to switch.&lt;/p&gt; 
&lt;h2&gt;What consumption pricing actually rewards&lt;/h2&gt; 
&lt;p&gt;Consumption pricing rewards customers for being efficient and using less. The customer is incentivized to run as efficiently as possible. But, that means the vendor is rewarded for inefficiency. A company billing by the query, the node hour, or the gigabyte makes more money when your workloads run slower, scan more data, or need more compute. Every optimization you make to your own pipeline is lost revenue for the vendor.&lt;/p&gt; 
&lt;p&gt;Clusters that sit idle between jobs still bill for the reserved time. Queries that scan an entire table instead of a filtered partition cost more, not less. The incentive isn't for the vendor to make the product faster. It's for the vendor to leave the meter running.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;Every optimization you make to your own pipeline is lost revenue for the vendor.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;h2&gt;When metering reflects real cost, and when it doesn't&lt;/h2&gt; 
&lt;p&gt;Consumption pricing makes sense for some products because the underlying cost actually is a function of usage. A water utility charges by the gallon because every gallon has to be pulled from a source, treated, and pushed to your kitchen sink. Use more, and the utility genuinely spends more. The meter reflects a real cost.&lt;/p&gt; 
&lt;p&gt;That's not how most data platforms work. When a customer runs a query against a consumption-priced data platform, the vendor isn't incurring a new marginal cost that scales the way a utility's does. The underlying cloud compute does cost more at higher volumes, but that's largely a pass-through cost the vendor marks up, not one they are absorbing on the customer's behalf. The relationship between treating a gallon of water versus a hundred gallons is direct, while the increased cost of processing a gigabyte and processing a petabyte on the same platform does not rise as directly. The actual cost structure does not require consumption to be profitable. The gallon of water is close to a true variable cost. The data billing model borrows from utilities but is applied to a product where the vendor's own cost structure doesn't actually require it.&lt;/p&gt; 
&lt;p&gt;That's the part of consumption pricing that doesn't hold up under scrutiny. It borrows the language and logic of utility billing, pay for what you use, without the underlying economics that make utility billing fair in the first place. For a lot of vendors, consumption pricing isn't reflecting their costs. It's a convenient way to capture more revenue as a customer's usage grows, dressed up as fairness.&lt;/p&gt; 
&lt;h2&gt;Why we priced Falcon differently&lt;/h2&gt; 
&lt;p&gt;Falcon customers pay a fixed annual fee based on the maximum number of concurrent cores used, starting at 64 cores for tier 1. Another advantage to our pricing is that it easily separates our licensing from your infrastructure decisions. Falcon can be run in a variety of conditions — on premises, cloud, air gapped, or edge.&lt;/p&gt; 
&lt;p&gt;We charge a fixed annual fee by throughput tier because we wanted our incentives pointed at the thing customers actually care about: Falcon running fast and efficiently on their data. If we billed by usage, every improvement we made to query performance would shrink our own revenue. Under fixed pricing, we get paid the same either way. Our customers want it fast, so we make it fast. Our customers want it efficient, so we make it efficient. That's the incentive we'd rather build a company around.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;Our customers want it efficient, so we make it efficient. That's the incentive we'd rather build a company around.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;Customers like that. They also like the predictability. A consumption bill can be filled with surprises. At the speed of today's businesses, usage can skyrocket. They are hard to forecast: Data volume grows, a new team starts running queries against the same cluster, a job that used to run once a day starts running every hour. A fixed annual fee based on capacity means a customer knows the cost of running Falcon before they sign, not after a surprise invoice.&lt;/p&gt; 
&lt;h2&gt;What this means for a customer&lt;/h2&gt; 
&lt;p&gt;Don't get us wrong; consumption-based pricing is not bad faith or fraudulent. We just wanted to build our company on the best incentives, and the tiered model works best for us. Fixed pricing points us toward the outcomes that customers really want. A customer evaluating Falcon against a consumption-priced alternative isn't just comparing two numbers (although we do reduce our overhead costs pretty significantly). They're comparing what each vendor is incentivized to optimize for.&lt;/p&gt; 
&lt;p&gt;If you're currently on a consumption-priced platform and want to know what a fixed-price alternative would actually cost, the &lt;a href="https://blog.haevek.com/falcon-test-flight" style="color: #f26b1d;"&gt;Falcon Test Flight&lt;/a&gt; runs your workloads against Falcon on your own data before any commercial conversation.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 32px 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Get started&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;&lt;a href="https://haevek.com/contact" style="color: #f26b1d; font-weight: 600;"&gt;Contact Haevek&lt;/a&gt; to learn more or sign up for a test flight.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-top: 3px solid #F26B1D; padding-top: 24px; margin: 40px 0 0 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Keep reading&lt;/p&gt; 
 &lt;p style="margin: 0 0 12px 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/efficiency-consumption-pricing" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Why Consumption Pricing Doesn't Reward Efficiency&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;The vendor incentive problem in detail, and how Falcon's capacity tiers pass efficiency gains to the customer.&lt;/span&gt;&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/spark-is-free-running-it-isnt" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Spark Is Free. Running It Isn't.&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why the free license and the infrastructure bill are two different conversations.&lt;/span&gt;&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Fno-consumption-pricing&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Cost Optimization</category>
      <category>Falcon Compute Platform</category>
      <pubDate>Mon, 21 Sep 2026 15:15:37 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/no-consumption-pricing</guid>
      <dc:date>2026-09-21T15:15:37Z</dc:date>
    </item>
    <item>
      <title>Uncovering the Real Cost of Serverless Infrastructure at Scale</title>
      <link>https://blog.haevek.com/posts/real-cost-of-serverless-at-scale</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Lowering the per-query cost &lt;em&gt;before&lt;/em&gt; you reach the valley of death.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Lowering the per-query cost &lt;em&gt;before&lt;/em&gt; you reach the valley of death.&lt;/p&gt;  
&lt;p&gt;Serverless infrastructure appeals to cost-conscious teams moving from prototype to early production, but kills their effectiveness once the application scales. The worst consequence&amp;nbsp;is not in the bills, it is how the pricing structure disincentivizes innovation.&lt;/p&gt; 
&lt;p&gt;Every major cloud provider and managed data platform has built a serverless offering that promises no cluster management, e.g. no cold start delays from spinning up Apache Spark instances, and no always-on costs for a workload that only runs a handful of queries a day. These offerings create a cost-effective methodology to get your project off the ground.&lt;/p&gt; 
&lt;p&gt;But the cost-effectiveness of serverless changes as you move to production-level query volumes. What felt like a reasonable price to pay at low query frequency turns into unwieldy fixed overhead as frequency climbs. And the architecture choices that looked like operational simplicity early on start producing operating costs that don't line up with the actual work being done. It has a chilling effect on work; useful queries and projects are held back from resource and budget constraints.&lt;/p&gt; 
&lt;h2&gt;Why Serverless Works at Low Query Volume&lt;/h2&gt; 
&lt;p&gt;The core problem serverless solves is real. A traditional Apache Spark cluster takes minutes to start from cold. This makes always-on infrastructure a practical necessity for any application with a latency requirement. Databricks' serverless SQL warehouses and equivalents from other vendors keep resource pools ready, scale to handle bursts automatically, and charge based on what you use rather than what you reserve. For a team running ten queries a day on a new product, paying a premium to skip cluster management and cold start delays saves both time and money.&lt;/p&gt; 
&lt;p&gt;The management overhead it removes is just as real. Cluster configuration, version upgrades, capacity planning, and operational monitoring all consume engineering time that a small team building a product often can't spare. Paying a managed service to absorb that burden, in exchange for a pricing premium, is a fair exchange when engineering hours are the scarcer resource.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;For most products at production-level query volume, serverless infrastructure costs more in dollars than it saves in engineering resources.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;h2&gt;Serverless Does Not Offer Economies of Scale&lt;/h2&gt; 
&lt;p&gt;To state the obvious, the higher operations and maintenance cost embedded in per-use serverless pricing rises along with query volume. A workload that has grown from ten queries per day to a thousand is paying the same overhead per unit as it did at low volume, meaning your costs have risen by a hundredfold. The benefit that overhead buys (cold start avoidance, burst handling, platform management) delivers less marginal value because the workload is now continuous enough to justify dedicated infrastructure.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px 32px 28px 32px; margin: 32px 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Why Falcon is different&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6; color: #d6d9de;"&gt;The patent-pending Falcon Compute Engine is inherently more efficient, not just at ramp up. Using&amp;nbsp;compiled code instead of the Java Virtual Machine and built Kubernetes-native, Falcon regularly cuts compute requirements by 90 percent and operations bills by 80 percent compared to Databricks alone. Kubernetes handles the scaling orchestration in a way that is familiar to dev teams as opposed to the proprietary serverless solutions of SaaS and infrastructure vendors.&lt;/p&gt; 
&lt;/div&gt; 
&lt;p&gt;The result is a &lt;em&gt;"valley of death"&lt;/em&gt; for data applications. You've developed a working prototype that should be a success story at scale, but the serverless infrastructure creates untenable per unit costs. Serverless got you this far, but now it is an obstacle to profitability.&lt;/p&gt; 
&lt;h2&gt;Maximum Throughput Made Cost Effective&lt;/h2&gt; 
&lt;p&gt;Falcon's cold start time for production workloads runs in seconds, matching the responsiveness profile of a serverless warehouse. The architecture, built on Kubernetes, is different; a small, fixed infrastructure footprint handles baseline query load, scales up on demand for burst requests, and scales back down once demand eases. This creates the cold start performance of serverless without the consumption pricing that makes serverless expensive at volume.&lt;/p&gt; 
&lt;div style="border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 28px 32px;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Deploying the Haevek Falcon Platform in the Real World&lt;/p&gt; 
 &lt;p style="margin: 0 0 16px 0; line-height: 1.6;"&gt;A customer using Databricks Serverless SQL Warehouses for end-user-facing queries reached the point where the serverless premium exceeded the benefit. The query volume was continuous enough that dedicated infrastructure made more sense, but managing a traditional Apache Spark cluster for a lean team was not a realistic alternative.&lt;/p&gt; 
 &lt;p style="margin: 0 0 16px 0; line-height: 1.6;"&gt;A pricing pattern in Databricks workloads compounds this effect. When a portion of a batch job benefits from the Photon engine, the Photon rate applies to the full job runtime, not only to the operations Photon accelerates. A job where 20 percent of the work runs on Photon gets billed at Photon rates for the full workload. At high job frequency, that pricing model produces infrastructure costs that significantly exceed what the compute throughput would suggest.&lt;/p&gt; 
 &lt;p style="margin: 0 0 16px 0; line-height: 1.6;"&gt;After deploying Falcon, the same query workload ran on approximately &lt;strong&gt;one-third of the prior infrastructure footprint&lt;/strong&gt; at approximately &lt;strong&gt;one-tenth of the operating cost.&lt;/strong&gt; The cold start response time matched what users had experienced with the serverless warehouse. The infrastructure scaled to query load during business hours and scaled back overnight.&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;The same configuration flexibility applies across multiple use cases. For applications where query latency is the priority, Falcon allocates the infrastructure for maximum throughput. For batch workloads where cost efficiency matters more than speed, the same engine runs at maximum savings. The configuration changes; the workload code does not.&lt;/p&gt; 
&lt;/div&gt; 
&lt;h2&gt;Lower Infrastructure Costs Unlock Your Full Potential&lt;/h2&gt; 
&lt;p&gt;Saving money is nice but elevating your capabilities is better. The Haevek customers who replaced serverless infrastructure with Falcon did not just use the savings to process the same workload at lower cost; they processed more. The analytical scope that was previously too expensive to run continuously became affordable. Queries that ran daily for cost reasons ran hourly. Reports that sampled data for budget reasons ran against full populations.&lt;/p&gt; 
&lt;p&gt;Economists call this the &lt;a href="https://en.wikipedia.org/wiki/Jevons_paradox" style="color: #f26b1d;"&gt;Jevons Paradox&lt;/a&gt;: as a resource becomes cheaper to use, consumption tends to increase because lower costs make previously unaffordable applications viable. Lower infrastructure cost does not mean teams will run fewer queries. It means they will run the queries they have been unable to justify running.&lt;/p&gt; 
&lt;p&gt;If serverless infrastructure costs are a constraint on what your application can afford to do at production scale, the &lt;a href="https://haevek.com/contact" style="color: #f26b1d;"&gt;Falcon test flight&lt;/a&gt; runs on your actual query workloads in your own environment before any commercial commitment.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 32px 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Get started&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;&lt;a href="https://haevek.com/contact" style="color: #f26b1d; font-weight: 600;"&gt;Contact Haevek&lt;/a&gt; to learn more or sign up for a test flight.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-top: 3px solid #F26B1D; padding-top: 24px; margin: 40px 0 0 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Keep reading&lt;/p&gt; 
 &lt;p style="margin: 0 0 12px 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/efficiency-consumption-pricing" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Why Efficiency and Consumption Pricing Don't Mix&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why consumption pricing works against the customer at production scale, and how Falcon's capacity model is built differently.&lt;/span&gt;&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/spark-is-free-running-it-isnt" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Spark Is Free. Running It Isn't.&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why the free license and the infrastructure bill are two different conversations.&lt;/span&gt;&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Freal-cost-of-serverless-at-scale&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Apache Spark</category>
      <category>Cost Optimization</category>
      <category>Falcon Compute Platform</category>
      <category>Databricks</category>
      <pubDate>Wed, 12 Aug 2026 21:40:00 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/real-cost-of-serverless-at-scale</guid>
      <dc:date>2026-08-12T21:40:00Z</dc:date>
    </item>
    <item>
      <title>Don't Take Our Word for It. Test It on Your Own Data.</title>
      <link>https://blog.haevek.com/posts/falcon-test-flight</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;A structured engagement that benchmarks Falcon against your actual production workloads — on your data, in your architecture, before any commercial commitment.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;A structured engagement that benchmarks Falcon against your actual production workloads — on your data, in your architecture, before any commercial commitment.&lt;/p&gt;  
&lt;p&gt;A test flight is a structured engagement where Haevek's customer engineering team configures Falcon against your actual production workloads, collects empirical benchmarks, and delivers a value assessment covering speed, cost, and infrastructure reduction on your specific architecture. The engineering work takes up to two weeks. The total engagement runs two to eight weeks depending on your organization's complexity. The test flight costs you nothing. If Falcon meets the agreed benchmarks, both parties move directly into commercial negotiations.&lt;/p&gt; 
&lt;p&gt;The benchmarks &lt;a href="https://haevek.com/use-cases" style="color: #f26b1d;"&gt;Haevek has published on this site&lt;/a&gt; came from real production deployments and will hold up to independent scrutiny. The business case for adopting Falcon should rest on your data, your architecture, and your infrastructure cost profile. A test flight produces that.&lt;/p&gt; 
&lt;h2&gt;Is this the right next step?&lt;/h2&gt; 
&lt;p&gt;A test flight runs against a subset of your pipelines, not your full stack. The goal is to generate a validated proof point on your specific architecture: empirical data that tells you whether the performance and cost improvement is real enough to justify rolling Falcon out across the rest of your workloads.&lt;/p&gt; 
&lt;p&gt;The workloads worth choosing for the test are the ones representative of your broader job portfolio. A batch ETL pipeline, a streaming workload, and an analytics or inference job cover the patterns most teams run. A result on those three gives you a credible basis for estimating what the full program looks like.&lt;/p&gt; 
&lt;p&gt;If you can recognize your situation in any of the &lt;a href="https://haevek.com/use-cases" style="color: #f26b1d;"&gt;use case results published here&lt;/a&gt;, the test flight is how you find out whether those numbers hold on your data.&lt;/p&gt; 
&lt;h2&gt;What you bring&lt;/h2&gt; 
&lt;p&gt;Two things are required from you: a defined use case with the pipelines you want to test, and reference data plus a current benchmark.&lt;/p&gt; 
&lt;p&gt;The use case definition does not have to be elaborate. A concrete pipeline description is enough: an upstream data source, the transformations applied, and the downstream destination. A batch ingest or ETL workload, a batch analytics or AI inference workload, and a streaming ETL workload represent the three most common starting configurations, and any combination of up to three pipelines works. If you have source code, bring it. Source code makes the validation process clean: Haevek's team can confirm the Falcon implementation performs exactly the same transformations at every step. Pseudocode and detailed process descriptions work as well and require additional validation effort to confirm equivalence.&lt;/p&gt; 
&lt;p&gt;For the benchmark, bring whatever measurement you have. How long does the pipeline run, on what infrastructure, at what cost? Many teams do not track this precisely for batch jobs that run overnight. If you do not have it measured, collecting a clean baseline becomes part of week one. The before/after comparison is the centerpiece of the case study report, so the baseline gets measured properly regardless of whether it arrives at kickoff or gets established during setup.&lt;/p&gt; 
&lt;p&gt;Haevek can run the test in its own environment or in yours depending on your data sensitivity and deployment requirements. The engagement is designed to run largely on Haevek's side. Your team joins for kickoff, provides access to data and the current benchmark, and checks in at agreed points. The lift is light.&lt;/p&gt; 
&lt;h2&gt;How the engagement runs&lt;/h2&gt; 
&lt;p&gt;Week one is setup. Haevek's team runs technical interchange meetings to understand the current state of your workloads, produces design documentation covering what functional equivalence looks like for each pipeline, and deploys into your environment if your security requirements call for it. Any compliance or security accreditation processes that need to run alongside the technical work start here, so there is no delay between the end of the test flight and a production deployment.&lt;/p&gt; 
&lt;p&gt;The following weeks are configuration and evaluation running in parallel. Haevek builds three variants of each workload. The first wraps your existing logic and runs it through Falcon's orchestration layer. The second pulls specific transformations into Rust using more of the Falcon standard library. The third is a full refactor of the pipeline built natively on Falcon.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.15em; line-height: 1.5; font-weight: 600; color: #0b0c0e;"&gt;Those three variants show you the performance improvement available at different levels of migration investment. Minimal adoption produces a meaningful result. Full adoption produces the headline performance numbers from the use case benchmarks.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;Having all three data points lets you estimate the migration effort for your team and what you get back at each stage.&lt;/p&gt; 
&lt;p&gt;Training is available throughout the engagement: your engineers can go through structured sessions on how to deploy, manage, and build new workloads on the Falcon Compute Platform. Teams that want to co-develop the refactored variants alongside Haevek's engineers can do that too.&lt;/p&gt; 
&lt;h2&gt;What you get&lt;/h2&gt; 
&lt;p&gt;The engagement closes with a case study report covering performance characteristics, scaling criteria, known limitations, and a value assessment of the configured Falcon workloads. Depending on what is most actionable for your team, that takes the form of a written report, a slide deck, or a working session.&lt;/p&gt; 
&lt;p&gt;Haevek's standard test flight measures against three success criteria, with additional metrics established per the specific use case and customer environment:&lt;/p&gt; 
&lt;ul&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Performance improvement:&lt;/strong&gt; a demonstrable improvement in end-to-end software runtime on equivalent or lower-cost infrastructure, benchmarked against your prior environment.&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;TCO reduction:&lt;/strong&gt; a material reduction in total cost of ownership (software licensing, compute infrastructure, maintenance, and support) over a twelve-month post-implementation period, aligned to outcomes agreed at the start of the engagement and benchmarked against the prior environment.&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Scalability:&lt;/strong&gt; maintained or improved performance and cost efficiency as workload size or concurrency increases, demonstrated through multi-executor or multi-node scaling tests.&lt;/li&gt; 
&lt;/ul&gt; 
&lt;p&gt;Haevek's customer engineering team has met all three criteria in every test flight completed to date. When the results come in, you have empirical data on your own workloads. If the criteria are met, the path to production is immediate. Haevek's security and compliance processes run in parallel from week one, so there is no gap between the end of the test flight and a production deployment. The team can move straight to making impact.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 32px 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Book your test flight&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;&lt;a href="https://haevek.com/contact" style="color: #f26b1d; font-weight: 600;"&gt;Contact Haevek&lt;/a&gt; to learn more or sign up for a test flight.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-top: 3px solid #F26B1D; padding-top: 24px; margin: 40px 0 0 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Keep reading&lt;/p&gt; 
 &lt;p style="margin: 0 0 12px 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/databricks-falcon-faq" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Replacing Databricks Compute with Falcon: Frequently Asked Questions&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Unity Catalog, Photon, Kubernetes support, benchmark access scope, and what stays on Databricks.&lt;/span&gt;&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Slash Your Databricks Compute Bill by over 80 Percent Without Migrating Anything&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why 70–80% of your bill sits in production compute, and how to find your offload candidates.&lt;/span&gt;&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Ffalcon-test-flight&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Apache Spark</category>
      <category>Falcon Compute Platform</category>
      <category>Databricks</category>
      <pubDate>Wed, 12 Aug 2026 21:39:49 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/falcon-test-flight</guid>
      <dc:date>2026-08-12T21:39:49Z</dc:date>
    </item>
    <item>
      <title>Why Consumption Pricing Doesn't Reward Efficiency</title>
      <link>https://blog.haevek.com/posts/efficiency-consumption-pricing</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Consumption pricing is cost-effective for prototyping, not scalability. That's why we built Falcon the way we did.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Consumption pricing is cost-effective for prototyping, not scalability. That's why we built Falcon the way we did.&lt;/p&gt;  
&lt;p&gt;&lt;img src="https://blog.haevek.com/hs-fs/hubfs/Consumption%20pricing%20header.png?width=936&amp;amp;height=526&amp;amp;name=Consumption%20pricing%20header.png" width="936" height="526" alt="Consumption pricing header" style="height: auto; max-width: 100%; width: 936px;"&gt;&lt;/p&gt; 
&lt;p&gt;If you are managing enterprise technology stacks, you no doubt are familiar with SaaS, IaaS, and other consumption-based pricing. It can be liberating for some scenarios and a prison in others. We built and license Falcon to break free of the traps in consumption-based pricing — and even the hidden cost traps of open source projects. With consumption pricing, the vendor is incentivized to increase customer usage regardless of efficiency. With capacity pricing, Haevek is incentivized to free up customer resources for innovation.&lt;/p&gt; 
&lt;p&gt;Consumption pricing is well-suited for early-stage work. You pay for what you use, run experiments when you need to, and stop paying when you stop running. For a team in prototyping mode, that model is genuinely useful, because costs stay proportional to actual activity and you don't pay for what you don't need.&lt;/p&gt; 
&lt;p&gt;That advantage disappears once workloads go to production and are always on. The ability to dial the meter up and down, which made consumption pricing the right choice during development, is gone. What remains is a recurring cost that increases with data volume and has no guaranteed ceiling.&lt;/p&gt; 
&lt;p&gt;Consider a data team with a $1 million annual infrastructure budget that can affordably prototype a new analytics pipeline for $10,000. At production scale, that same pipeline may run $200,000 a month. By May, the budget is gone.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; line-height: 1.6; color: #0b0c0e;"&gt;Uber's CTO described &lt;a href="https://www.theinformation.com/newsletters/applied-ai/uber-cto-shows-claude-code-can-blow-ai-budgets" style="color: #f26b1d;"&gt;exactly this dynamic&lt;/a&gt; to The Information earlier this year after the company burned through its entire annual AI tooling budget in four months: "The budget I thought I would need is blown away already."&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;The cloud compute costs that were controllable at low usage become impossible to predict when adoption runs at production scale. Maintaining workloads requires far more budget than planned. Continuing to run at scale eats through money but shutting it down disrupts production users who have come to depend on it.&lt;/p&gt; 
&lt;p&gt;&lt;img src="https://blog.haevek.com/hs-fs/hubfs/2026-08%20Why%20Efficiency%20and%20Consumption%20Pricing%20Dont%20Mix-graphic-1.png?width=2544&amp;amp;height=1633&amp;amp;name=2026-08%20Why%20Efficiency%20and%20Consumption%20Pricing%20Dont%20Mix-graphic-1.png" width="2544" height="1633" alt="2026-08 Why Efficiency and Consumption Pricing Dont Mix-graphic-1" style="height: auto; max-width: 100%; width: 2544px;"&gt;&lt;/p&gt; 
&lt;p style="font-size: 0.9em; color: #666666; margin: 8px 0 32px 0;"&gt;The flexibility that consumption pricing offered in development costs transforms into overhead that grows as you scale up.&lt;/p&gt; 
&lt;h2&gt;The Vendor Incentive Problem&lt;/h2&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;At production scale, the incentives built into consumption pricing create a windfall for the vendor and work against the people paying for it.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;Under consumption pricing, vendors earn less revenue from the customers who create and benefit from efficiency improvements. If their software runs twice as fast, the customer pays half as much in compute time. That is a direct hit to the vendor's revenue.&lt;/p&gt; 
&lt;p&gt;A vendor in that position has two options. Accept the lower revenue or recapture the loss through pricing adjustments, pushing the aggregate costs back up to approximately where they were before the gain.&lt;/p&gt; 
&lt;p&gt;This is a structural property of the consumption pricing model, like a team billing by the hour earning more by working slower. Cloud compute and consumption-based data platforms work the same way. When a customer needs more throughput, the vendor's incentive points them toward adding more compute. When a customer needs to cut the bill, every optimization they make reduces what they pay the vendor.&lt;/p&gt; 
&lt;p&gt;Haevek has watched this play out on programs where Spark was the processing framework. Engineering teams had already optimized those pipelines as far as the model allowed. Configurations were tuned, code was clean, and costs increased in step with growing data volume. Every available cost saving tool within that pricing structure had reached its limit. The only remaining options were to spend more or find an approach that ran the same workloads at lower cost.&lt;/p&gt; 
&lt;p style="font-weight: bold;"&gt;Firsthand experience like this is the primary reason Falcon's pricing model is built the way it is.&lt;/p&gt; 
&lt;h2&gt;A Pricing Model That Rewards Efficiency&lt;/h2&gt; 
&lt;p&gt;Falcon's pricing model is built around capacity rather than consumption. A reasonable analogy is broadband internet. You buy a bandwidth tier, not billed per byte transferred.&lt;/p&gt; 
&lt;p&gt;Falcon licenses work on the same principle. You commit to an annual tier of concurrent compute cores. That covers everything you run on Falcon for the year. When Falcon ships performance improvements that make pipelines run faster, your licensed capacity handles more throughput, which means additional workloads fit within the same license at no additional cost.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;Under the Falcon pricing model, the benefit from performance improvement flows to the customer automatically.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;As Falcon's efficiency improves, it reduces cloud infrastructure consumption without affecting Falcon license costs. They complete the same work in less time, using fewer cloud resources, and the infrastructure savings stay with them.&lt;/p&gt; 
&lt;p&gt;Our pricing model is predictable, making it easier for our customers to plan and innovate. The infrastructure budget for the year is a known number from day one. Data volume growth does not translate into bill surprises. Our customers can confidently experiment and build knowing that there will not be unexpected costs caused by growth in data volume.&lt;/p&gt; 
&lt;h2&gt;What the Falcon Pricing Model Looks Like in Practice&lt;/h2&gt; 
&lt;p&gt;A data analytics company was spending several million dollars a year on Databricks and AWS. They were ingesting multiple terabytes of sensor data per day, transforming it, running analytics, and providing insights to their customers. Their infrastructure cost was growing at the same time their customers were asking for more complex analytics capabilities, and the existing workloads were consuming the budget that would have funded those capabilities.&lt;/p&gt; 
&lt;p&gt;&lt;img src="https://blog.haevek.com/hs-fs/hubfs/2026-08%20Why%20Efficiency%20and%20Consumption%20Pricing%20Dont%20Mix-graphic-2.png?width=2544&amp;amp;height=1687&amp;amp;name=2026-08%20Why%20Efficiency%20and%20Consumption%20Pricing%20Dont%20Mix-graphic-2.png" width="2544" height="1687" alt="2026-08 Why Efficiency and Consumption Pricing Dont Mix-graphic-2" style="height: auto; max-width: 100%; width: 2544px;"&gt;&lt;/p&gt; 
&lt;p&gt;The migration started with batch and streaming workloads. On batch processing, Falcon showed 80% cost reduction against the equivalent Databricks workflow. On streaming, Databricks was running 26 servers carrying both Databricks fees and the underlying AWS infrastructure costs. Two servers running Falcon have replaced those 26 servers, cutting AWS EC2 costs by 90%.&lt;/p&gt; 
&lt;p&gt;Eight days after signing the contract, they had their first pipeline in production and were seeing the savings accumulate. Within 90 days, they had expanded their license by 4x, finding more workloads where the same economics applied.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;The savings in those 90 days more than covered the annual cost of the Falcon license.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;Databricks is still in their stack. It handles data science workloads, notebook environments, and dashboarding. End users who rely on those capabilities see no change. The underlying Delta Lake tables are unchanged. What changed is where the batch and streaming compute runs, and what that costs.&lt;/p&gt; 
&lt;p&gt;They escaped the consumption pricing trap; new feature investment no longer demanded steep budget commitments for consumption-based pricing before there was revenue. The freed up infrastructure budget went to developing robust analytics capabilities their customers were asking for.&lt;/p&gt; 
&lt;div style="border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 28px 32px;"&gt; 
 &lt;p style="margin: 0 0 16px 0; line-height: 1.6;"&gt;Those paying a Databricks bill are familiar with the Databricks unit (DBU). The DBU is the per-second/minute/hour pricing for Databricks. The total cost of Databricks is DBUs plus cloud infrastructure costs. If 80% of that DBU allocation is currently going to batch and streaming workloads, moving those to Falcon means that 80% of the committed spend can go toward the features Databricks genuinely excels at:&lt;/p&gt; 
 &lt;ul style="margin: 0 0 16px 0;"&gt; 
  &lt;li style="margin-bottom: 8px;"&gt;data science tooling,&lt;/li&gt; 
  &lt;li style="margin-bottom: 8px;"&gt;notebooks,&lt;/li&gt; 
  &lt;li style="margin-bottom: 8px;"&gt;dashboards,&lt;/li&gt; 
  &lt;li style="margin-bottom: 8px;"&gt;and the newer AI and analytics capabilities.&lt;/li&gt; 
 &lt;/ul&gt; 
 &lt;p style="margin: 0; line-height: 1.6; font-weight: 600;"&gt;The total cost of operations stays the same. Capabilities expand substantially.&lt;/p&gt; 
&lt;/div&gt; 
&lt;h2&gt;Comparing Falcon, Databricks, and Spark Costs&lt;/h2&gt; 
&lt;p&gt;Falcon is built to be the most efficient distributed compute engine possible so businesses can be more strategic and innovative. Consumption pricing does not support our goal. Under consumption pricing, every efficiency improvement Falcon shipped would reduce Haevek's own revenue. Our pricing model, based on cores running our Kubernetes based Falcon platform, had to reflect what the engine is actually built to do.&lt;/p&gt; 
&lt;p&gt;The other challenge was competing against open source software, and free is a difficult baseline for a pricing conversation. Apache Spark is free to download. The infrastructure it runs on is not, and often charged on a consumption pricing model.&lt;/p&gt; 
&lt;p&gt;Falcon is on average &lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d;"&gt;80% faster than Spark on equivalent workloads&lt;/a&gt;. This translates into needing 80% less cloud infrastructure to do the same work, which, like the Databricks scenario above, saves more than the cost of the Falcon license. The license tiers are structured so that 50% to 80% of those infrastructure savings stay with the customer after the Falcon license cost is covered, with larger tiers returning a higher share of the savings.&lt;/p&gt; 
&lt;h2&gt;Finding the Right Tier with the Falcon Test Flight&lt;/h2&gt; 
&lt;p&gt;Before any license commitment, the Falcon test flight runs on the customer's actual workloads and collects benchmarks. From that data, Haevek can model a full-scale infrastructure scenario; from 10% of the pipeline we project the full workload and determine the best fit Falcon license tier. A team that needs to justify a budget decision can choose based on measured performance against their own data, rather than a projection built on assumptions.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;Haevek wants the customer's first dollar to be backed by evidence and be on a pricing model that rewards both parties for running workloads as efficiently as possible.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;The &lt;a href="https://blog.haevek.com/falcon-test-flight" style="color: #f26b1d;"&gt;Falcon test flight&lt;/a&gt; runs your actual workloads on your own infrastructure so you can see what those numbers look like before any commitment is made. If you're carrying the infrastructure cost of a Spark-based production workload and want to know whether our pricing model works for you, &lt;a href="https://haevek.com/contact" style="color: #f26b1d;"&gt;get started here&lt;/a&gt;.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 32px 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Get started&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;&lt;a href="https://haevek.com/contact" style="color: #f26b1d; font-weight: 600;"&gt;Contact Haevek&lt;/a&gt; to learn more or sign up for a test flight.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-top: 3px solid #F26B1D; padding-top: 24px; margin: 40px 0 0 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Keep reading&lt;/p&gt; 
 &lt;p style="margin: 0 0 12px 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/spark-is-free-running-it-isnt" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Spark Is Free. Running It Isn't.&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why the free license and the infrastructure bill are two different conversations.&lt;/span&gt;&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/why-good-data-projects-die" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Why Good Data Projects Die Before They Ship&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Nine out of ten proven prototypes never reach production. The killer is the economics, not the technology.&lt;/span&gt;&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Fefficiency-consumption-pricing&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Cost Optimization</category>
      <category>Falcon Compute Platform</category>
      <category>Databricks</category>
      <pubDate>Wed, 12 Aug 2026 21:39:23 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/efficiency-consumption-pricing</guid>
      <dc:date>2026-08-12T21:39:23Z</dc:date>
    </item>
    <item>
      <title>Spark Is Free. Running It Isn't.</title>
      <link>https://blog.haevek.com/posts/spark-is-free-running-it-isnt</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;The license is free. The compute bill that follows it is not — and at production scale, it sets the ceiling on what your team can build.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;The license is free. The compute bill that follows it is not — and at production scale, it sets the ceiling on what your team can build.&lt;/p&gt;  
&lt;p&gt;Downloading Apache Spark costs nothing, which shapes how most data teams frame their infrastructure budget. The software is free, so the cloud bill gets treated as a cloud cost, disconnected from the software choice driving the compute consumption. But the cloud bill is a direct function of how efficiently the compute layer runs, and Spark's compute efficiency has a ceiling that most teams hit well before they run out of analytical ambition.&lt;/p&gt; 
&lt;p&gt;At production data volumes (terabytes or more processed daily, pipelines running continuously, workloads that cannot pause when the budget runs short), compute becomes a dominant cost in a data team's operations. A streaming workload at that scale typically runs on hundreds of cores and costs hundreds of thousands to millions of dollars per year. The Spark license is free, and the infrastructure bill that accumulates alongside it is substantial.&lt;/p&gt; 
&lt;p&gt;&lt;a href="https://haevek.com/use-cases" style="color: #f26b1d;"&gt;Benchmarks across a wide range of production workloads&lt;/a&gt; show Falcon delivering an average 80 percent reduction in infrastructure consumption against equivalent Spark deployments. That reduction generates savings large enough to cover the Falcon license and leave the customer around 40 percent better off on total spend, making the paid software cheaper to operate than the free version.&lt;/p&gt; 
&lt;h2&gt;The compute cost survives the managed service decision&lt;/h2&gt; 
&lt;p&gt;Every team running Spark at production scale pays for compute. The choice between self-hosting and using Databricks or Amazon EMR affects who manages the operational complexity, leaving the underlying compute bill unchanged.&lt;/p&gt; 
&lt;p&gt;Self-hosting Spark means paying directly for cloud VMs or on-premises hardware, plus the engineering overhead of managing the platform. Databricks and EMR add a management premium on top of that infrastructure spend in exchange for handling cluster configuration, platform updates, and operational complexity. The compute bill runs in both cases, and the managed service charges for the work of managing it.&lt;/p&gt; 
&lt;p&gt;The relative merits of Databricks against EMR against self-managed Spark are a legitimate operational question, and the answer varies by team size and technical capability. The compute cost itself is a separate and often larger variable. An engine that runs the same workloads in one-fifth the infrastructure changes the economics of either option.&lt;/p&gt; 
&lt;h2&gt;What Spark costs at production scale&lt;/h2&gt; 
&lt;p&gt;A real-time streaming workload processing multiple terabytes per day on Databricks runs on 432 cores and costs $464,000 per year in infrastructure. That reflects the cost profile of a well-run production Spark workload at that data volume.&lt;/p&gt; 
&lt;p&gt;The cost compounds as pipeline complexity grows. Adding analytical steps, inline models, or embeddings generation pushes workloads further into compute-bound territory, where Spark's JVM-based runtime and garbage collection overhead extract a higher infrastructure price for each unit of useful output. A workload that runs cheaply at sample scale becomes expensive at population scale because the volume exposes what the runtime costs to operate.&lt;/p&gt; 
&lt;p&gt;Haevek has encountered government intelligence programs operating on fixed monthly infrastructure budgets where the compute allocation runs out before the month ends. Analytical workloads shut down and mission owners wait for the next billing cycle. A data analytics pipeline processing multiple terabytes of sensor data daily was capped at city-scale analysis because country-scale processing exceeded what the available compute budget could support.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;In both cases the capability constraint was financial. The infrastructure bill set the analytical ceiling, and the analytical ceiling determined what questions the team could ask.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;h2&gt;The same workloads at one-fifth the cost&lt;/h2&gt; 
&lt;p&gt;&lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d;"&gt;Falcon runs the same Spark workloads at around 80 percent lower infrastructure cost on average&lt;/a&gt;. The streaming workload that cost $464,000 per year on 432 cores runs on Falcon for $66,000 per year on 32 cores. The batch ETL workload that ran on 26 Databricks servers runs on two servers running Falcon, with EC2 costs down over 90 percent. On Falcon, healthcare OCR and LLM inference pipelines run 75 percent faster at 80 percent lower infrastructure cost, and anomaly detection engines processing 200 million events show 95 percent runtime reduction at 95 percent lower cost.&lt;/p&gt; 
&lt;p&gt;&lt;a href="https://blog.haevek.com/efficiency-consumption-pricing" style="color: #f26b1d;"&gt;Falcon's pricing model&lt;/a&gt; is structured to pass those savings to the customer. License tiers are priced on a capacity basis: infrastructure consumption drops by approximately 80 percent, and after the Falcon license cost is covered, roughly 40 percent of prior total spend stays with the customer as net savings. On a $100,000 monthly Spark infrastructure bill, the new total (infrastructure plus Falcon license) comes to roughly $60,000. The software carries a cost, and the total spend is lower than running the free version.&lt;/p&gt; 
&lt;h2&gt;What lower compute cost makes possible&lt;/h2&gt; 
&lt;p&gt;Eight days after signing a Falcon contract, the first pipeline was in production. Within 90 days, the company had expanded its license four times as more workloads moved across, and the infrastructure savings generated in that period exceeded the full annual license cost.&lt;/p&gt; 
&lt;p&gt;The city-scale analytical ceiling lifted as the compute cost dropped, and country-scale processing became a standard part of operations.&lt;/p&gt; 
&lt;p&gt;The pattern follows what economists call Jevons Paradox. As infrastructure becomes cheaper to use, teams use more of it, because lower costs bring previously unaffordable work into reach. The analyses scoped down to fit the budget get run at full scale, and the questions that previously exceeded the cost threshold get asked.&lt;/p&gt; 
&lt;p&gt;If Spark or Databricks infrastructure spend is a constraint on what your team can build or what questions you can answer, the &lt;a href="https://blog.haevek.com/falcon-test-flight" style="color: #f26b1d;"&gt;Falcon test flight&lt;/a&gt; runs on your actual workloads in your own environment before any commercial commitment to help you understand how much compute cost you can save with Falcon.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 32px 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Get started&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;&lt;a href="https://haevek.com/contact" style="color: #f26b1d; font-weight: 600;"&gt;Contact Haevek&lt;/a&gt; to learn more or sign up for a test flight.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-top: 3px solid #F26B1D; padding-top: 24px; margin: 40px 0 0 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Keep reading&lt;/p&gt; 
 &lt;p style="margin: 0 0 12px 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/efficiency-consumption-pricing" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Why Efficiency and Consumption Pricing Don't Mix&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why consumption pricing works against the customer at production scale, and how Falcon's capacity model is built differently.&lt;/span&gt;&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Slash Your Databricks Compute Bill by over 80 Percent Without Migrating Anything&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why 70–80% of your bill sits in production compute, and how to find your offload candidates.&lt;/span&gt;&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Fspark-is-free-running-it-isnt&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Apache Spark</category>
      <category>Cost Optimization</category>
      <category>Falcon Compute Platform</category>
      <pubDate>Wed, 12 Aug 2026 21:38:31 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/spark-is-free-running-it-isnt</guid>
      <dc:date>2026-08-12T21:38:31Z</dc:date>
    </item>
    <item>
      <title>Nobody Would Be Stupid Enough to Rebuild Apache Spark (So We Did)</title>
      <link>https://blog.haevek.com/posts/nobody-would-rebuild-apache-spark</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;The founding story of the Falcon engine: two failed versions, a language decision made over a winter holiday, and the architectural conflict that took three rewrites to resolve.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;The founding story of the Falcon engine: two failed versions, a language decision made over a winter holiday, and the architectural conflict that took three rewrites to resolve.&lt;/p&gt;  
&lt;p&gt;&lt;span style="white-space-collapse: preserve;"&gt; &lt;img src="https://blog.haevek.com/hs-fs/hubfs/undefined-1.png?width=1672&amp;amp;height=941&amp;amp;name=undefined-1.png" width="1672" height="941" alt="undefined-1" style="height: auto; max-width: 100%; width: 1672px;"&gt;&lt;/span&gt;&lt;/p&gt; 
&lt;p&gt;There's an article arguing that no startup would be stupid enough to rebuild Apache Spark from scratch. The reasoning is sound: a decade of development history, thousands of contributors, and a scope that would exhaust any startup's runway well before producing anything useful.&lt;/p&gt; 
&lt;p&gt;&lt;span style="font-weight: bold;"&gt;We found that article while we already had a working version.&lt;/span&gt;&lt;/p&gt; 
&lt;p&gt;Getting there took longer than we expected and required solving problems we hadn't anticipated. This is the story of what those were.&lt;/p&gt; 
&lt;p&gt;The founding problem surfaced twice, from different directions. The first was a government program running a 15-year-old distributed compute system written in C++ — fast enough to meet processing requirements, but so complex that it needed several dozen engineers with PhDs and two decades of C++ experience just to keep it running. Apache Spark seemed like the obvious path to a leaner architecture. The engineers on the program had already tried it. It wasn't fast enough. The second was a rare event prediction program where dev environments alone were running $100,000 to $200,000 a month in compute before anything reached production. The customer was direct: you can't keep throwing infrastructure at this problem. Both situations pointed to the same conclusion: Spark needed to be faster, and the only way to get there was to rebuild the execution engine itself.&lt;/p&gt; 
&lt;p&gt;That conclusion was easy to reach and uncomfortable to sit with. Rebuilding it in C++ was the obvious move technically, but C++ puts the entire burden of memory safety on the developer. The pattern is well documented across production systems: memory mismanagement accounts for the majority of exploitable vulnerabilities in software built in C++, and the maintenance cost follows. A team that had spent years cleaning up poorly written C++ code, tracing memory leaks with unreliable tooling and rewriting implementations that worked on paper but couldn't be sustained in production, had no interest in building that problem into a new foundation. We spent a winter holiday benchmarking several major backend systems languages available: C, C++, Go, Python, Rust, and Java. Rust matched or exceeded C++-level performance across every test that mattered. It enforces memory safety at compile time through strict ownership and borrowing rules, with no garbage collector, which means no GC pauses and no runtime overhead. Go was close on performance but came with a garbage collector, which rules it out for production compute workloads requiring consistent throughput. The choice was straightforward once the data was in front of us.&lt;/p&gt; 
&lt;p&gt;NSA and CISA reached the same conclusion independently. Their &lt;a href="https://media.defense.gov/2025/Jun/23/2003742198/-1/-1/0/CSI_MEMORY_SAFE_LANGUAGES_REDUCING_VULNERABILITIES_IN_MODERN_SOFTWARE_DEVELOPMENT.PDF" style="color: #f26b1d;"&gt;joint guidance published in June 2025&lt;/a&gt; recommends memory-safe languages as a primary strategy for reducing software vulnerabilities. Google Project Zero found that 75% of CVEs exploited in the real world were memory safety vulnerabilities. Android reduced its share from 76% to 24% by switching new development to Rust. For Haevek's Department of Defense, Intelligence Community, and NATO customers, building in Rust meets both a performance requirement and a compliance mandate their own agencies are pushing the industry toward.&lt;/p&gt; 
&lt;p&gt;The deployment architecture came from a separate insight: a piece on Medium arguing that the JVM was dead and the new virtual machine was Linux in a container on Kubernetes. Years of deploying large-scale software into air-gapped government environments, getting systems to bootstrap without remote access and work through authority-to-operate processes for classified deployments at an enterprise AI platform company, had made clear what Kubernetes could do. If compiled Rust code in a container carried essentially no overhead, the same compute engine could run in a hyperscaler, in a customer data center, or on an edge node with no internet connection. We validated that overhead assumption the same winter holiday. It was unmeasurable at the scales we were operating at.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;Then we started building the engine, and the first two versions failed.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;The problem that drove both failures, and that took a third attempt to resolve, is a fundamental architectural conflict. High-performance compiled languages require static typing: you define at compile time how the whole system is going to work, and that's where the performance comes from. But a general-purpose distributed compute engine has to handle any workload, with different data types, different pipeline shapes, different compute profiles, and different source and destination formats. Those two requirements pull directly against each other, and working out how to satisfy both simultaneously while keeping the system usable for developers who would build their jobs on top of it is what three full rewrites produced. The techniques developed to resolve that conflict are now the foundation of Haevek's patent portfolio, and they represent years of work that any team attempting this independently would have to produce on their own.&lt;/p&gt; 
&lt;p&gt;The Bauplan team made the same assessment in their &lt;a href="https://arxiv.org/pdf/2308.05368" style="color: #f26b1d;"&gt;2023 paper on building a serverless Data Lakehouse&lt;/a&gt;, describing a from-scratch rebuild as "a challenge unfit for a startup" and choosing to reuse existing components instead. For most infrastructure problems, that's the right approach. For what we were trying to do, Spark's performance ceiling isn't addressable through configuration or workload-level optimization. It's architectural. By the time we found the article saying nobody would attempt this, we'd already had a working version for some time.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px 32px 28px 32px; margin: 32px 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;A dozen person-years later&lt;/p&gt; 
 &lt;p style="margin: 0 0 20px 0; font-size: 1.35em; line-height: 1.4; font-weight: 600; color: #ffffff;"&gt;Falcon processes the same distributed workloads as Spark at up to 25 times faster performance, and deploys anywhere Kubernetes runs.&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6; color: #d6d9de;"&gt;The engine already covers approximately 70 to 80 percent of the functionality data engineering teams use day to day. The first customers to put Falcon into production replaced their Databricks workloads in eight days.&lt;/p&gt; 
&lt;/div&gt; 
&lt;p&gt;In one case, running against compiled code written for a government program, Falcon measured up to six times faster performance on certain operations. Two factors likely explain this: Falcon's resource allocation model uses Kubernetes to go well beyond the standard one-thread-per-CPU assumption and aligns compute footprint to whether a workload is network-bound, disk-bound, or compute-bound. Developers working in compiled languages tend to optimize the inner loop while treating serialization, deserialization, queuing, and cache management as details the language will handle. Those aren't small inefficiencies at scale, and significant engineering effort has gone into each of them precisely because that's where performance gets lost when nobody's paying attention.&lt;/p&gt; 
&lt;p&gt;If you want to see what the numbers look like on your workload rather than ours, that's what &lt;a href="https://haevek.com/contact" style="color: #f26b1d;"&gt;the Falcon test flight&lt;/a&gt; is for. The Falcon test flight takes a real workload from your environment, runs it side by side against your existing implementation on equivalent infrastructure, and delivers the result before any commercial commitment. The claims in this article were earned the same way.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 32px 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Get started&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;&lt;a href="https://haevek.com/contact" style="color: #f26b1d; font-weight: 600;"&gt;Contact Haevek&lt;/a&gt; to learn more or sign up for a test flight.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-top: 3px solid #F26B1D; padding-top: 24px; margin: 40px 0 0 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Keep reading&lt;/p&gt; 
 &lt;p style="margin: 0 0 12px 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Slash Your Databricks Compute Bill by over 80 Percent Without Migrating Anything&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why 70–80% of your bill sits in production compute, and how to find your offload candidates.&lt;/span&gt;&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/spark-is-free-running-it-isnt" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Spark Is Free. Running It Isn't.&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why the free license and the infrastructure bill are two different conversations.&lt;/span&gt;&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Fnobody-would-rebuild-apache-spark&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Apache Spark</category>
      <category>Falcon Compute Platform</category>
      <category>Data Engineering</category>
      <pubDate>Wed, 12 Aug 2026 20:37:20 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/nobody-would-rebuild-apache-spark</guid>
      <dc:date>2026-08-12T20:37:20Z</dc:date>
    </item>
    <item>
      <title>Why Good Data Projects Die Before They Ship</title>
      <link>https://blog.haevek.com/posts/why-good-data-projects-die</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;The prototype works. Everyone agrees it should ship. Then the production cost analysis kills it — and the pattern repeats across government and commercial programs alike.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;The prototype works. Everyone agrees it should ship. Then the production cost analysis kills it — and the pattern repeats across government and commercial programs alike.&lt;/p&gt;  
&lt;p&gt;&lt;span style="white-space-collapse: preserve;"&gt; &lt;img src="https://blog.haevek.com/hs-fs/hubfs/Good%20data%20projects.webp?width=1672&amp;amp;height=941&amp;amp;name=Good%20data%20projects.webp" width="1672" height="941" alt="Good data projects" style="height: auto; max-width: 100%; width: 1672px;"&gt;&lt;/span&gt;&lt;/p&gt; 
&lt;p&gt;The idea works perfectly in a notebook. A data scientist builds a classification model, runs it against a sample dataset, and the results look exactly like what the program owner asked for. The technique is sound, the output is demonstrably useful, and everyone in the room agrees this should go into production. Across government and commercial programs, roughly nine out of ten projects that reach this point never make it through.&lt;/p&gt; 
&lt;blockquote&gt; 
 &lt;p style="font-weight: bold;"&gt;&lt;span style="font-size: 24px;"&gt;The failure is almost never the underlying idea. What kills these projects is the difference between what something costs to prove and what it costs to run at the scale the business actually requires.&lt;/span&gt;&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;h2&gt;The noise problem&lt;/h2&gt; 
&lt;p&gt;Take sentiment analysis on a handful of customer support transcripts. Running it on a sample is cheap, fast, and the results look compelling. Applied to every support conversation a large consumer business handles in a day, the same process becomes prohibitively expensive, not because the technique is flawed but because at production scale the vast majority of the data being processed carries no meaningful signal. In a support chat context, somewhere between 90 and 99 percent of conversations are routine exchanges: package tracking, password resets, standard billing questions. The compute cycles spent determining that someone asked where their order is are wasted cycles, and at volume they compound into serious infrastructure spend.&lt;/p&gt; 
&lt;p&gt;This is the pattern that recurs across data engineering contexts: a prototype gets validated on a representative sample, and the sample is necessarily information-dense because that is what makes a good demonstration. Production data is not information-dense. Most of it is noise. When the production environment floods an expensive model or pipeline with background chatter, efficiency collapses and cost explodes.&lt;/p&gt; 
&lt;p&gt;The more tractable approach is to treat expensive compute as a scarce resource rather than a general-purpose tool. A system that efficiently identifies the ten seconds of genuine signal within ten minutes of noise and routes only that to downstream processing will dramatically outperform one that runs everything through the expensive layer. The technique is sound, and the architecture around it determines whether the economics work.&lt;/p&gt; 
&lt;h2&gt;Why it keeps happening&lt;/h2&gt; 
&lt;p style="font-weight: bold;"&gt;Several forces compound to produce this pattern with regularity.&lt;/p&gt; 
&lt;p&gt;At the executive level, business needs get translated into vague mandates and passed to engineering teams without the full cost context or a clear definition of what success looks like at production volume. A program owner asks for fraud detection in financial transactions, or sentiment identification across customer support, or anomaly detection in sensor telemetry. The engineer builds something that meets the stated technical objective, often without any analysis of whether that objective is worth the cost of meeting it at the scale the business actually generates. The second-order question of whether the system will be sustainable never surfaces until the infrastructure bill does. The pattern showed up recently at a financial services organization that needed to add a feature to share data with customers in a specific compliant format. Rather than evaluating off-the-shelf tooling that could have addressed the need in two to three weeks, the default response was to assign four engineers and give them a year to build a custom solution. The result was an order of magnitude more in engineering time spent and a feature that reached customers ten times later than it needed to.&lt;/p&gt; 
&lt;p&gt;There is also significant organizational pressure across most large enterprises to demonstrate that AI is being applied to business problems. Completing an AI initiative has become a success metric in its own right, independent of whether the initiative produces value that outweighs its cost. This creates a dynamic where prototypes get built, demos get presented, and the decision to go to production gets driven by enthusiasm rather than economics.&lt;/p&gt; 
&lt;p&gt;Engineering bias compounds both of those forces. Data engineers and data scientists are drawn to technically interesting problems. Building a complex LLM-powered pipeline is more compelling than implementing&amp;nbsp;a rule-based classifier. The more novel the approach, the more organizational momentum it gathers, regardless of whether the complexity is necessary for the problem being solved. By the time the infrastructure cost becomes apparent, significant time and budget have already been committed.&lt;/p&gt; 
&lt;h2&gt;The math that kills the project&lt;/h2&gt; 
&lt;p&gt;The economics of this problem look different depending on whether the return on investment is easy to define. The outcome is similar across all three cases.&lt;/p&gt; 
&lt;p&gt;When ROI is indirect and hard to measure, a budget downturn exposes the weakness immediately. A large multinational conglomerate built three supply chain applications covering risk management, inventory management, and supply visibility, which aggregated data from dozens of sources into dashboards that gave teams genuinely faster access to the information they needed. Users liked the product. Specific use cases demonstrated that data which had previously taken weeks to pull together was now available in seconds. The problem was that the value delivered was indirect: better decisions, easier jobs, faster access to information. When business conditions tightened and the infrastructure cost came under scrutiny, the value was difficult to defend because it did not show up as a line item saving. After years of investment, the organization attempted to rebuild the capability using internally developed tooling to reduce costs. That helped but did not solve the problem. The infrastructure required to normalize, enrich, and cache data from that many sources at the required frequency still could not be justified, and the program was significantly scaled back.&lt;/p&gt; 
&lt;p&gt;When ROI is clear and directly measurable, the math still often does not work. County governments have a well-understood case for property appraisal tooling: counties routinely lag the market rate on property valuations, sometimes by years, and more accurate and timely appraisals translate directly into increased tax revenue. The value of getting this right is quantifiable before a single line of code is written. The challenge is that an accurate appraisal system requires a large volume of localized data and continuous model retraining as market conditions change. The compute cost of training cycles and inference pipelines, combined with the frequency at which models need to be updated to stay accurate, makes the economics viable only for large counties with substantial technology budgets already in place. Across state and local government programs, dozens of property appraisal projects have been started, demonstrated good results in testing, and then been cancelled at the cost analysis stage.&lt;/p&gt; 
&lt;p&gt;When ROI is difficult to define at all, the infrastructure requirement still has to be justified against something. Law enforcement intelligence applications are a clear example: the goal is to make it easier to investigate crimes and identify connections across data from multiple agencies. The value of improving investigative capacity is real but nearly impossible to convert into a number. The infrastructure requirement is extreme, with data from multiple sources that is frequently incomplete or incorrect requiring heavy ETL pipelines to correct and infer, combined with a hard requirement for fast query response times that forces large datasets to be cached in graph structures. Those graphs are expensive to build and expensive to maintain. Across programs in this space, the cost analysis has cancelled more projects than any technical obstacle has.&lt;/p&gt; 
&lt;h2&gt;What changes when the cost does&lt;/h2&gt; 
&lt;p&gt;The question worth asking before building is not whether something is technically possible. It almost always is. The question is whether it is economically viable at the scale and the user base the business actually has.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;Innovation programs that consistently produce promising prototypes with no path to production are not failing because the ideas are bad. They are failing because the economics of getting those ideas to production and keeping them there do not work.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;When the infrastructure cost of running distributed compute workloads drops substantially, more projects cross the viability threshold. The supply chain dashboards that users genuinely valued become fundable. The property appraisal models that worked in testing become deployable for mid-sized counties, not just the largest ones. The law enforcement intelligence applications that had to be abandoned after successful pilots get a second chance at production.&lt;/p&gt; 
&lt;p&gt;That is the infrastructure cost problem Falcon is built to solve. By processing distributed workloads at substantially lower cost than Apache Spark, it changes the economics of the production step that kills most data projects before they ship. If you have workloads that proved themselves in testing but stalled at the cost analysis, the &lt;a href="https://haevek.com/contact" style="color: #f26b1d;"&gt;Falcon test flight&lt;/a&gt; runs your actual jobs on your own infrastructure so you can see what those numbers look like before any commercial commitment.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 32px 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Get started&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;&lt;a href="https://haevek.com/contact" style="color: #f26b1d; font-weight: 600;"&gt;Contact Haevek&lt;/a&gt; to learn more or sign up for a test flight.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-top: 3px solid #F26B1D; padding-top: 24px; margin: 40px 0 0 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Keep reading&lt;/p&gt; 
 &lt;p style="margin: 0 0 12px 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/efficiency-consumption-pricing" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Why Efficiency and Consumption Pricing Don't Mix&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why consumption pricing works against the customer at production scale, and how Falcon's capacity model is built differently.&lt;/span&gt;&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/falcon-test-flight" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Don't Take Our Word for It. Test It on Your Own Data.&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;What a test flight involves, what you bring, and the success criteria it measures against.&lt;/span&gt;&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Fwhy-good-data-projects-die&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Cost Optimization</category>
      <category>Falcon Compute Platform</category>
      <category>Data Engineering</category>
      <pubDate>Wed, 12 Aug 2026 20:25:52 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/why-good-data-projects-die</guid>
      <dc:date>2026-08-12T20:25:52Z</dc:date>
    </item>
    <item>
      <title>Replacing Databricks Compute with Falcon: Frequently Asked Questions</title>
      <link>https://blog.haevek.com/posts/databricks-falcon-faq</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Straight answers to the questions data engineering and platform teams ask before moving production compute off the JVM.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Straight answers to the questions data engineering and platform teams ask before moving production compute off the JVM.&lt;/p&gt;  
&lt;p&gt;Most teams evaluating Haevek's Falcon Compute Platform arrive with the same set of questions: what breaks, what stays, what access we need, and how long it takes to find out. Below are the answers, drawn from customer benchmarks and production deployments.&lt;/p&gt; 
&lt;p&gt;For the full argument behind these answers, see &lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d;"&gt;Slash Your Databricks Compute Bill by over 80 Percent Without Migrating Anything&lt;/a&gt;. For the operational sequence, see &lt;a href="https://blog.haevek.com/databricks-offload-playbook" style="color: #f26b1d;"&gt;The Databricks Offload Playbook&lt;/a&gt;.&lt;/p&gt; 
&lt;h2&gt;Platform and Architecture&lt;/h2&gt; 
&lt;h3&gt;What is the Falcon Compute Platform?&lt;/h3&gt; 
&lt;p&gt;Falcon is a distributed compute engine built in Rust as a complete ground-up redesign of Apache Spark. Every data processing step compiles to machine instructions that execute directly on the CPU, with optimal offloading to specialized chips (GPUs, TPUs, and similar) should workloads require it. It runs Kubernetes-native: each pipeline executes as an atomic K8s job, so compute activates when the job starts and releases when it finishes.&lt;/p&gt; 
&lt;h3&gt;Does replacing Databricks compute require migrating data or changing Unity Catalog configurations?&lt;/h3&gt; 
&lt;p&gt;No. Falcon is a pure compute engine that reads from and writes to your existing Delta Lake or Iceberg tables in your current object store. Unity Catalog remains the governance layer with no re-registration, no access policy changes, and no table contract modifications. Your lakehouse storage layer is completely untouched.&lt;/p&gt; 
&lt;h3&gt;Which Kubernetes environments does Falcon support?&lt;/h3&gt; 
&lt;p&gt;Falcon deploys on any flavor of k8s including: EKS (AWS), AKS (Azure), GKE (Google Cloud), and OpenShift. It installs into your existing Kubernetes environment inside your own cloud account or on-premise environment — no new accounts or storage layers are required.&lt;/p&gt; 
&lt;h3&gt;How does Falcon handle AI inference workloads alongside batch ETL?&lt;/h3&gt; 
&lt;p&gt;Falcon supports direct bindings to existing libraries across languages, so teams can call Python, C/C++, or other dependencies natively within a pipeline without a separate service layer. As one example of what this enables, ONNX (Open Neural Network Exchange, an open standard format for packaging and sharing machine learning models) models can be embedded directly in the pipeline, eliminating the separate model-serving infrastructure many teams run alongside Databricks for batch inference. This reduces both operational complexity and the number of compute footprints to manage.&lt;/p&gt; 
&lt;h2&gt;Cost and the JVM&lt;/h2&gt; 
&lt;h3&gt;Why don't standard Databricks optimizations like Photon or auto-termination reduce the bill enough?&lt;/h3&gt; 
&lt;p&gt;Photon can reduce job runtime for certain analytical queries, but Databricks charges substantially more DBUs for Photon-enabled compute, making the cost impact largely net neutral. It also doesn't support every operation a pipeline might use, so many jobs only partially benefit from it even when it is enabled. Auto-termination only affects idle interactive clusters, which represent 20–30% of a typical enterprise bill. The 70–80% sitting in "always-on" production compute is structurally unaffected by either optimization.&lt;/p&gt; 
&lt;h3&gt;What is the JVM overprovisioning multiplier and why does it matter for Databricks cost?&lt;/h3&gt; 
&lt;p&gt;&lt;span style="color: #5c4f6e; background-color: #fafaf7;"&gt;The JVM interprets Java bytecode, meaning every operation passes through an abstraction layer rather than compiling directly to machine instructions. When a job processes terabytes of data, the inefficiency of the abstraction layer has measurable impacts. For example, it constantly serializes records into JVM objects, ships those objects between workers across the network, and deserializes them on the other side.&amp;nbsp;&lt;/span&gt;&lt;span style="color: #5c4f6e; background-color: #fafaf7;"&gt;Another impact is that the &lt;/span&gt;JVM's garbage collector periodically halts all threads to reclaim memory — a behavior known as stop-the-world GC. Since these pauses are unpredictable, teams typically provision significantly more compute than a workload theoretically requires just to reliably hit SLA windows. The result is a structural cost floor that cluster tuning and right-sizing can reduce at the margins but can't eliminate. Switching to a compiled native runtime removes that floor entirely.&lt;/p&gt; 
&lt;h3&gt;Does switching to a managed alternative like EMR or Dataproc solve this?&lt;/h3&gt; 
&lt;p&gt;No. Whether you run a managed service (Databricks, Amazon EMR, Azure HDInsight, Google Dataproc) or self-managed Spark, the underlying cloud compute bill stays the same. What varies is what you pay on top: managed services add a licensing layer, while self-managed Spark trades that licensing cost for the engineering hours needed to maintain and tune the cluster yourself. None of those choices reduce the compute footprint itself, because the JVM inefficiency runs underneath all of them.&lt;/p&gt; 
&lt;h3&gt;How much can a team realistically expect to save?&lt;/h3&gt; 
&lt;p&gt;It depends on how much of your bill is production compute and how heavily you're overprovisioned. For a representative mid-market account with a $400K annual Databricks bill where 75% is production compute, that is $300K running on the JVM that can be migrated. Applying the 85% TCO reduction observed in the commercial streaming benchmark — a figure that already accounts for Falcon's license cost — yields roughly $255K in annualized savings, with the remaining $100K of interactive work staying on Databricks untouched. Teams running 5–10x excess capacity see the largest absolute savings, because the efficiency gain applies to the biggest base.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px 32px 28px 32px; margin: 32px 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Benchmark · Commercial data analytics company&lt;/p&gt; 
 &lt;p style="margin: 0 0 20px 0; font-size: 1.35em; line-height: 1.4; font-weight: 600; color: #ffffff;"&gt;432 cores on Databricks. 32 cores on Falcon. A 93% reduction in compute infrastructure.&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6; color: #d6d9de;"&gt;A production real-time streaming ETL workload processing multi-terabytes per day. Same transforms, same data. Annual spend dropped from $464K to $66K for this job, an 85% reduction in total cost of ownership.&lt;/p&gt; 
&lt;/div&gt; 
&lt;h2&gt;Evaluation and Scope&lt;/h2&gt; 
&lt;h3&gt;How long does a Falcon benchmark take, and what access does it require?&lt;/h3&gt; 
&lt;p&gt;The benchmark window runs 2–4 weeks for the side-by-side measurement phase. Falcon needs narrowly scoped read access to specific source object store paths and write access to a designated output path. No access to your Databricks workspace or Unity Catalog governance layer is required.&lt;/p&gt; 
&lt;h3&gt;Which workloads should stay on Databricks?&lt;/h3&gt; 
&lt;p&gt;Interactive SQL exploration, notebook-based ML experimentation, Delta Live Tables pipelines tightly coupled to Databricks orchestration, and anything where analyst iteration speed matters more than infrastructure efficiency. The goal is not to leave Databricks. It is to shift the 70–80% of spend running on production compute off the JVM runtime and keep Databricks where it earns its place.&lt;/p&gt; 
&lt;h3&gt;How do I identify which of my jobs to move first?&lt;/h3&gt; 
&lt;p&gt;Query &lt;code style="background: #F2F3F5; padding: 2px 6px; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.9em;"&gt;system.billing.usage&lt;/code&gt; over a 30-day window and sort descending by total DBUs. Flag anything running more than four hours per execution, or anything with a run count near 720 for the month — that's an hourly job, which means it is always on. Three categories are typically worth targeting: long-running batch ETL, 24/7 streaming ingestion, and high-throughput analytics or ML inference. The full query is in &lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d;"&gt;the main article&lt;/a&gt;.&lt;/p&gt; 
&lt;h3&gt;How is the cutover sequenced so we don't take an outage?&lt;/h3&gt; 
&lt;p&gt;Shadow mode first: Falcon runs alongside the existing pipeline on the same input with no production traffic routed to it, so you can compare outputs and benchmark cost before anything shifts. Once output parity and performance targets are confirmed, you route production traffic to Falcon and keep the rollback path live until you are satisfied that it can be shut down. For API-driven workloads, an optional canary phase shifts 5–10% of traffic incrementally until all traffic is migrated over to Falcon. The full sequence is in &lt;a href="https://blog.haevek.com/databricks-offload-playbook" style="color: #f26b1d;"&gt;The Databricks Offload Playbook&lt;/a&gt;.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 40px 0 0 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Question not answered here?&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;Email &lt;a href="mailto:info@haevek.com?subject=Test%20Flight" style="color: #f26b1d; font-weight: 600;"&gt;info@haevek.com&lt;/a&gt; with the subject line "Test Flight."&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Fdatabricks-falcon-faq&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Apache Spark</category>
      <category>Cost Optimization</category>
      <category>Falcon Compute Platform</category>
      <category>Databricks</category>
      <pubDate>Wed, 12 Aug 2026 17:56:51 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/databricks-falcon-faq</guid>
      <dc:date>2026-08-12T17:56:51Z</dc:date>
    </item>
    <item>
      <title>The Databricks Offload Playbook: Cut Over Without Downtime</title>
      <link>https://blog.haevek.com/posts/databricks-offload-playbook</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;The operational sequence for moving a live batch or streaming pipeline off Databricks compute and onto Haevek's Falcon Compute Platform, without an outage.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;The operational sequence for moving a live batch or streaming pipeline off Databricks compute and onto Haevek's Falcon Compute Platform, without an outage.&lt;/p&gt; 
&lt;p&gt;Deciding to move production compute off the JVM is the easy part. The hard part is the sequencing: how you get a live, business-critical pipeline from "running on Databricks" to "running on Falcon" without an outage, without a correctness regression, and without a rollback story you have to invent under pressure.&lt;/p&gt; 
&lt;p&gt;This is the playbook. It assumes you've already worked out &lt;em&gt;why&lt;/em&gt; production compute is where your Databricks bill actually lives — if you haven't, start with &lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d;"&gt;Slash Your Databricks Compute Bill by over 80 Percent Without Migrating Anything&lt;/a&gt;.&lt;/p&gt; 
&lt;h2&gt;Before You Start: Have a Shortlist&lt;/h2&gt; 
&lt;p&gt;You need a specific workload, not a general intention. Run the &lt;code style="background: #F2F3F5; padding: 2px 6px; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.9em;"&gt;system.billing.usage&lt;/code&gt; query in &lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d;"&gt;the Prioritizing Workloads section of this article&lt;/a&gt; against a 30-day window, sort descending by total DBUs, and pull your top candidates. Three categories are typically worth targeting: long-running batch ETL, 24/7 streaming ingestion pipelines that never scale to zero, and high-throughput analytics or AI/ML inference jobs.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.15em; line-height: 1.5; font-weight: 600; color: #0b0c0e;"&gt;One thing does not change at any point in this sequence: your data stays where it is. Falcon reads from and writes to your existing Delta Lake or Iceberg tables, and Unity Catalog remains the governance layer.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;There is no data migration, no catalog re-registration, and no change to access policies or table contracts. What you are changing is the engine, not the lakehouse.&lt;/p&gt; 
&lt;h2&gt;Cutover Patterns&lt;/h2&gt; 
&lt;p&gt;Cutting over a live pipeline to a new compute layer without downtime requires sequencing. The patterns below lay out that sequence from first validation to full production.&lt;/p&gt; 
&lt;div style="border-left: 3px solid #F26B1D; padding: 4px 0 4px 24px; margin: 24px 0;"&gt; 
 &lt;p style="margin: 0 0 8px 0; font-size: 1.1em; font-weight: bold; color: #0b0c0e;"&gt;Shadow Mode&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;Run Falcon alongside your existing pipeline without routing any production traffic to it. Both systems process the same input so you can compare outputs, catch semantic differences early, and benchmark actual runtime and cost against your baseline before anything shifts.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-left: 3px solid #F26B1D; padding: 4px 0 4px 24px; margin: 24px 0;"&gt; 
 &lt;p style="margin: 0 0 8px 0; font-size: 1.1em; font-weight: bold; color: #0b0c0e;"&gt;Full Cutover&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;Once shadow mode confirms output parity and performance targets are met, route all production traffic to Falcon and decommission the old pipeline. Keep the rollback path live for a period of time&amp;nbsp;as a precaution.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-left: 3px solid #C9CDD3; padding: 4px 0 4px 24px; margin: 24px 0;"&gt; 
 &lt;p style="margin: 0 0 8px 0; font-size: 1.1em; font-weight: bold; color: #0b0c0e;"&gt;Optional — Canary and Parallel Validation&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;For API-driven workloads where incremental traffic shifting is possible, route a small slice of production traffic (5–10%) to Falcon while the rest continues through your existing stack. Monitor output consistency, latency, and error rates, then increase Falcon's share incrementally until you're confident to cut over fully. This step is less applicable to batch job workloads and can be skipped where shadow mode results are sufficient.&lt;/p&gt; 
&lt;/div&gt; 
&lt;p&gt;The Falcon Test Flight compresses this entire sequence into a structured 2–4 week engagement. Security and compliance reviews run from week one, so production deployment and cutover can follow immediately once the engagement closes.&lt;/p&gt; 
&lt;h2&gt;The Four Steps&lt;/h2&gt; 
&lt;h3&gt;Step 1 — Pick One Workload&lt;/h3&gt; 
&lt;p&gt;Pull the top three results from your &lt;code style="background: #F2F3F5; padding: 2px 6px; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.9em;"&gt;system.billing.usage&lt;/code&gt; query and pick the highest-spend batch or streaming job that isn't a single point of failure for a customer-facing pipeline. You want a meaningful savings signal and a real benchmark, not maximum risk on your first run.&amp;nbsp;&lt;/p&gt; 
&lt;h3&gt;Step 2 — Deploy To Production&lt;/h3&gt; 
&lt;p&gt;Haevek installs Falcon into your EKS, AKS, or GKE environment in your own cloud account or your variant of kubernetes in your on-premise environment. No new cloud accounts, no new storage layers, and no changes to Unity Catalog. The deployment uses narrowly scoped permissions: Falcon needs access to read your source data and write to a designated output path. Implementation across two pipelines has run in as few as 7 days to get into production.&lt;/p&gt; 
&lt;h3&gt;Step 3 — Run In Parallel&lt;/h3&gt; 
&lt;p&gt;Both the Databricks job and the Falcon job consume the same source data and write to separate output paths. At the end of each run, compare output parity: schema, row counts, and any business-logic assertions your test suite already covers. Collect infrastructure metrics via Prometheus and Grafana in parallel. Run in parallel long enough to give you enough samples to smooth out throughput variance and establish a reliable compute-cost comparison.&lt;/p&gt; 
&lt;h3&gt;Step 4 — Cut Over and Repeat&lt;/h3&gt; 
&lt;p&gt;Once output parity is confirmed and benchmark results are in, deprecate the Databricks job for that workload. Keep Databricks running for interactive SQL, notebook development, and Unity Catalog governance — those workloads stay where they belong. Then return to your &lt;code style="background: #F2F3F5; padding: 2px 6px; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.9em;"&gt;system.billing.usage&lt;/code&gt; shortlist and repeat the process for the next workload.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px 32px 28px 32px; margin: 32px 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;The compounding loop&lt;/p&gt; 
 &lt;p style="margin: 0 0 20px 0; font-size: 1.35em; line-height: 1.4; font-weight: 600; color: #ffffff;"&gt;Customers who've done this report expanding their Falcon footprint 4x within 90 days.&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6; color: #d6d9de;"&gt;Each successful migration unlocks the budget for the next and increases long term cost savings. The first production customer recovered their full annual license cost from TCO savings within 73 days of cutover.&lt;/p&gt; 
&lt;/div&gt; 
&lt;h2&gt;What Good Looks Like at Each Gate&lt;/h2&gt; 
&lt;p&gt;The measurement framework is the same one used during the Test Flight benchmark, and it should stay in place through cutover:&lt;/p&gt; 
&lt;ul&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Wall-clock runtime:&lt;/strong&gt; end-to-end job duration, not just execution time&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Compute cores consumed:&lt;/strong&gt; averaged across multiple runs, not a single sample&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Infrastructure cost at list pricing:&lt;/strong&gt; so the comparison is apples-to-apples regardless of your negotiated rates&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Output correctness:&lt;/strong&gt; schema parity, row counts, and any business-logic assertions your existing test suite covers&lt;/li&gt; 
&lt;/ul&gt; 
&lt;p&gt;Output correctness is the gate that decides whether you advance. Runtime and cost tell you whether the move was worth making; parity tells you whether you're allowed to make it. Do not collapse the two.&lt;/p&gt; 
&lt;h2&gt;Getting Started&lt;/h2&gt; 
&lt;p&gt;If you haven't run a benchmark yet, the Falcon Test Flight runs up to three of your production workloads side-by-side against your live Databricks jobs, in your environment, on your actual data, using your own infrastructure metrics. Success criteria are defined upfront and tailored to your specific workloads.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 32px 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Get started&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;Email &lt;a href="mailto:info@haevek.com?subject=Test%20Flight" style="color: #f26b1d; font-weight: 600;"&gt;info@haevek.com&lt;/a&gt; with the subject line "Test Flight."&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-top: 3px solid #F26B1D; padding-top: 24px; margin: 40px 0 0 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Keep reading&lt;/p&gt; 
 &lt;p style="margin: 0 0 12px 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/slash-databricks-compute-bill-without-migrating" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Slash Your Databricks Compute Bill by over 80 Percent Without Migrating Anything&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Why 70–80% of your bill sits in production compute, and what a 93% infrastructure reduction actually looks like.&lt;/span&gt;&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/databricks-falcon-faq" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Replacing Databricks Compute with Falcon: Frequently Asked Questions&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Unity Catalog, Photon, Kubernetes support, benchmark access scope, and what stays on Databricks.&lt;/span&gt;&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Fdatabricks-offload-playbook&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Falcon Compute Platform</category>
      <category>Databricks</category>
      <category>Data Engineering</category>
      <category>Migration</category>
      <pubDate>Wed, 12 Aug 2026 17:29:38 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/databricks-offload-playbook</guid>
      <dc:date>2026-08-12T17:29:38Z</dc:date>
    </item>
    <item>
      <title>Slash Your Databricks Compute Bill by over 80 Percent Without Migrating Your Data</title>
      <link>https://blog.haevek.com/posts/slash-databricks-compute-bill-without-migrating</link>
      <description>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Let Databricks do what Databricks does best.&amp;nbsp;Move batch and streaming tasks to Haevek's Falcon Compute Platform to save on the base load.&lt;/p&gt;</description>
      <content:encoded>&lt;p style="font-size: 1.25em; line-height: 1.5; font-style: italic; border-left: 4px solid #F26B1D; padding-left: 20px; margin: 0 0 32px 0;"&gt;Let Databricks do what Databricks does best.&amp;nbsp;Move batch and streaming tasks to Haevek's Falcon Compute Platform to save on the base load.&lt;/p&gt;  
&lt;p&gt;Your team has tried everything on the Databricks cost checklist. Auto-termination is running. Cluster policies are tuned. Batch jobs have been moved to Jobs Compute. You've evaluated Photon, reviewed query plans, and tightened idle cluster timeouts. And yet, at the end of every quarter, finance is still asking why the bill hasn't moved.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.3em; line-height: 1.45; font-weight: 600; color: #0b0c0e;"&gt;You can't right-size your way to 80% savings when the engine itself is the source of overprovisioning.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;The reason those optimizations haven't worked is that they're targeting the wrong part of the spend.&lt;/p&gt; 
&lt;h2&gt;The Limits of Standard Databricks Cost Advice&lt;/h2&gt; 
&lt;p&gt;Roughly &lt;a href="https://medium.com/@vishal.dutt.data.architect/cost-optimization-strategies-every-databricks-admin-should-know-0b5c2353ebe0#:~:text=compute%20is%2070%E2%80%9380%25%20of%20the%20bill" style="color: #f26b1d;"&gt;70–80% of a typical enterprise Databricks bill is production compute&lt;/a&gt;: batch ETL pipelines running nightly or hourly, and streaming ingestion jobs that run 24 hours a day, seven days a week. The remaining 20–30% is interactive work: analysts running ad-hoc queries, data scientists iterating in notebooks, exploratory SQL against Delta tables.&lt;/p&gt; 
&lt;p&gt;It used to be that running these costly streaming and batch processes on Databricks or Spark was the only path forward, despite its old scaling and resource approaches. To optimize, we tried to chip away at the small slice that is interactive work.&lt;/p&gt; 
&lt;p&gt;To understand why those production jobs are so expensive, you need to understand what Spark and the JVM are doing underneath Databricks.&lt;/p&gt; 
&lt;p&gt;&lt;a href="https://docs.databricks.com/aws/en/spark/faq" style="color: #f26b1d;"&gt;The heart of the Databricks platform is Apache Spark&lt;/a&gt;, which uses the Java Virtual Machine as its execution runtime. The JVM interprets Java bytecode, meaning every operation passes through an abstraction layer rather than compiling directly to machine instructions. When a job processes a terabyte of data, it constantly serializes records into JVM objects, ships those objects between workers across the network, and deserializes them on the other side. Between each of those steps, the garbage collector periodically pauses every thread while the JVM reclaims memory, adding latency that compounds across a large cluster.&lt;/p&gt; 
&lt;p&gt;To meet SLAs, it is necessary to provision a lot more compute capacity to account for garbage collection and the JVM abstraction layer — often 5–10x the footprint they theoretically require. If a pipeline needs 40 cores to process a daily ingestion volume within a given window, that same pipeline running on Spark typically requires 200–400 cores to reliably hit that window, because GC pauses, serialization overhead, and spike buffers all eat into the headroom.&lt;/p&gt; 
&lt;blockquote style="border: 0; border-left: 4px solid #F26B1D; background: #FDF4EE; margin: 32px 0; padding: 24px 28px;"&gt; 
 &lt;p style="margin: 0; font-size: 1.15em; line-height: 1.5; font-weight: 600; color: #0b0c0e;"&gt;If a pipeline needs 40 cores to process a daily ingestion volume within a given window, that same pipeline running on Spark typically requires 200–400 cores to reliably hit that window.&lt;/p&gt; 
&lt;/blockquote&gt; 
&lt;p&gt;Standard FinOps advice optimizes around the edges. You can switch to spot instances and you cut the per-hour rate, but you are still running 400 cores. Photon speeds certain analytical queries through vectorized execution outside the JVM (at a substantial cost premium), but any deviation from a Photon-supported function brings the JVM, GC overhead, and serialization costs back into play.&lt;/p&gt; 
&lt;p&gt;Regardless of whether you are running a managed service (e.g., Databricks, Amazon EMR, Azure HDInsight, Google Dataproc) or self-managed Spark, the underlying cloud compute bill still stays the same. What varies is what you pay on top of that: Databricks and equivalents add a licensing layer, while self-managed Spark trades that licensing cost for the engineering hours needed to maintain and tune the cluster yourself. Managed services exist largely because that operational and integration burden is real and significant. Yet none of these choices reduce the compute footprint itself, because the JVM inefficiencies and subsequent overprovisioning problem runs underneath all of them. You're still provisioning 5–10x more compute than you theoretically need, regardless of how you manage it.&lt;/p&gt; 
&lt;p style="font-size: 1.15em; font-weight: 600; color: #0b0c0e;"&gt;The only lever that changes the core-hours-per-unit-of-work ratio is replacing the compute engine itself.&lt;/p&gt; 
&lt;h2&gt;Why a 93% Infrastructure Reduction Is Architecturally Achievable&lt;/h2&gt; 
&lt;p&gt;Haevek's Falcon Compute Platform can take on those loads with less compute. It does not need the over provisioning as it is compiled into a native binary and does not require garbage collection.&lt;/p&gt; 
&lt;p&gt;Falcon is a distributed compute engine built in Rust as a complete ground-up redesign of Apache Spark. Every data processing step compiles to machine instructions that execute directly on the CPU with optimal offloading to specialized chips (GPUs, TPUs, etc.) when workloads require it. Falcon's pipeline model keeps data in contiguous memory structures rather than converting records to Java objects, so there's no memory overhead or serialization penalty between processing stages. The throughput-per-core is &lt;em&gt;structurally&lt;/em&gt; different from anything running on the JVM.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px 32px 28px 32px; margin: 32px 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Benchmark · Commercial data analytics company&lt;/p&gt; 
 &lt;p style="margin: 0 0 20px 0; font-size: 1.35em; line-height: 1.4; font-weight: 600; color: #ffffff;"&gt;432 cores on Databricks. 32 cores on Falcon. A 93% reduction in compute infrastructure.&lt;/p&gt; 
 &lt;p style="margin: 0 0 16px 0; line-height: 1.6; color: #d6d9de;"&gt;A production real-time streaming ETL workload was processing multi-terabytes per day on Databricks using 432 cores for a single pipeline. The same pipeline, running the same transforms on the same data, required 32 cores on Falcon. Annual spend dropped from $464K to $66K for this job, an 85% reduction in total cost of ownership (TCO).&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 0.9em; line-height: 1.6; color: #a8acb3;"&gt;The 93% figure reflects the compute footprint change. The 85% TCO figure is the &lt;em&gt;net&lt;/em&gt; comparison including infrastructure and software license costs for Falcon and Databricks.&lt;/p&gt; 
&lt;/div&gt; 
&lt;p&gt;Falcon also eliminates the always-on cluster problem through Kubernetes-native execution. Each pipeline runs as an atomic K8s job: compute activates when the job starts and releases when it finishes. For streaming workloads, Falcon bursts to meet throughput demand and scales back immediately, without maintaining a warm cluster between micro-batches. A Databricks streaming cluster bills continuously whether it's processing 100,000 events per second or 10.&lt;/p&gt; 
&lt;p&gt;You don't have to migrate data or re-architect; Falcon reads from and writes to the same lakehouse as Databricks. The new solution works with your existing object store's Delta Lake and Iceberg tables. Unity Catalog remains the governance layer, with no data migration, no catalog re-registration, and no changes to access policies or table contracts. Falcon is a pure compute engine and knows how to stay in its lane.&lt;/p&gt; 
&lt;p&gt;&lt;img src="https://blog.haevek.com/hs-fs/hubfs/Haevek-Falcon-Infographic%20(1).png?width=2525&amp;amp;height=2121&amp;amp;name=Haevek-Falcon-Infographic%20(1).png" width="2525" height="2121" alt="Haevek-Falcon-Infographic (1)" style="height: auto; max-width: 100%; width: 2525px;"&gt;&lt;/p&gt; 
&lt;p&gt;Falcon deploys on any Kubernetes environment such as EKS, AKS, GKE, OpenShift, or whichever is already running in your cloud account. Falcon supports direct bindings to existing libraries across languages, so teams can call Python, C/C++, or other dependencies natively within a single job without a separate service layer. For example, ONNX (Open Neural Network Exchange, an open standard format for packaging and sharing machine learning models) models can be embedded directly in the pipeline, eliminating the separate model-serving infrastructure many teams run alongside Databricks for batch inference.&lt;/p&gt; 
&lt;h2&gt;Prioritizing Workloads to Reduce Costs on the 80%&lt;/h2&gt; 
&lt;h3&gt;Building Your Offload Shortlist&lt;/h3&gt; 
&lt;p&gt;Right now, you can identify candidates for migration to Falcon. Before you change anything, you need to know exactly which jobs are costing you the most. Databricks System Tables make this straightforward.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Step 1.&lt;/strong&gt; Run the following query against &lt;code style="background: #F2F3F5; padding: 2px 6px; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.9em;"&gt;system.billing.usage&lt;/code&gt; filtered to a 30-day window:&lt;/p&gt; 
&lt;pre style="background: #0B0C0E; color: #d6d9de; padding: 24px; overflow-x: auto; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.85em; line-height: 1.6; margin: 24px 0;"&gt;&lt;code style="font-family: ui-monospace,SFMono-Regular,Menlo,monospace; color: #d6d9de; background: transparent;"&gt;-- Run this in your Databricks SQL workspace.
-- Ensure you have the appropriate permissions to query system tables
-- and that Unity Catalog is enabled in your workspace.

-- First, set the system catalog:
USE CATALOG system;

-- Verify your exact billing_origin_product values by running:
-- SELECT DISTINCT billing_origin_product FROM billing.usage LIMIT 100

SELECT
  COALESCE(usage_metadata.job_name, usage_metadata.job_run_id) AS job_identifier,
  billing_origin_product,
  sku_name,
  SUM(usage_quantity) AS total_dbus,
  COUNT(DISTINCT usage_metadata.job_run_id) AS run_count
FROM billing.usage
WHERE usage_date &amp;gt;= CURRENT_DATE() - INTERVAL 30 DAYS
  AND billing_origin_product IN ('JOBS', 'STREAMING', 'MODEL_SERVING')
GROUP BY
  COALESCE(usage_metadata.job_name, usage_metadata.job_run_id),
  billing_origin_product,
  sku_name
ORDER BY total_dbus DESC
LIMIT 20;&lt;/code&gt;&lt;/pre&gt; 
&lt;p&gt;&lt;strong&gt;Step 2.&lt;/strong&gt; Sort by &lt;code style="background: #F2F3F5; padding: 2px 6px; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.9em;"&gt;total_dbus&lt;/code&gt; descending. The top results are your offload candidates. Flag anything running more than four hours per execution or anything with a &lt;code style="background: #F2F3F5; padding: 2px 6px; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.9em;"&gt;run_count&lt;/code&gt; close to 720 for the month: that's a job running every hour, which means it's &lt;em&gt;always on&lt;/em&gt;.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Step 3.&lt;/strong&gt; Categorize items. Three categories are typically worth targeting:&lt;/p&gt; 
&lt;ul&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Long-running batch ETL:&lt;/strong&gt; Nightly or hourly transforms over large datasets, such as multi-terabyte ingestion, enrichment, and aggregation pipelines. These typically run on Jobs Compute and carry the JVM overhead multiplier across every execution.&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;24/7 streaming ingestion pipelines:&lt;/strong&gt; Always-on clusters that never scale to zero. A streaming cluster maintained for continuous processing bills continuously, even when throughput is low. That idle overhead is structural and can't be tuned away.&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;High-throughput analytics and AI/ML inference jobs:&lt;/strong&gt; Scheduled batch analytics, ML inference, data enrichment, or embedding generation pipelines. Think "what we run to make enriched gold data from silver data" in a medallion architecture. Spark was designed for table transforms, not matrix operations, which means these workloads tend to overprovision even more severely than standard ETL on the JVM.&lt;/li&gt; 
&lt;/ul&gt; 
&lt;p&gt;Some workloads belong on Databricks and should stay there: interactive SQL exploration, notebook-based ML experimentation, Delta Live Tables pipelines tightly coupled to Databricks orchestration, and anything where analyst iteration speed matters more than infrastructure efficiency. The goal is to shift the 70–80% of spend running on production compute off the JVM runtime and keep Databricks where it earns its place.&lt;/p&gt; 
&lt;h2&gt;The Test Flight: How the Side-by-Side Benchmark Actually Works&lt;/h2&gt; 
&lt;p&gt;Will Falcon work on these workloads? What kind of savings can you get? &lt;a href="https://haevek.com/contact" style="color: #f26b1d;"&gt;The Falcon Test Flight&lt;/a&gt; is designed to answer these questions without interfering with your operations or creating new costs.&lt;/p&gt; 
&lt;p&gt;Haevek engineers help you deploy Falcon into your cloud environment using narrowly scoped permissions: no access to your Databricks workspace, your governance layer, or any data plane beyond the specific object store paths needed for the benchmark workloads. You select up to three production-representative jobs — typically one batch ETL pipeline, one streaming ingestion job, and one inference workload if applicable — and run the same jobs in parallel with the same source data in both Falcon and Databricks, writing to separate output paths. The measurement framework covers four dimensions:&lt;/p&gt; 
&lt;ul&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Wall-clock runtime:&lt;/strong&gt; end-to-end job duration, not just execution time&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Compute cores consumed:&lt;/strong&gt; averaged across multiple runs, not a single sample&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Infrastructure cost at list pricing:&lt;/strong&gt; so the comparison is apples-to-apples regardless of your negotiated rates&lt;/li&gt; 
 &lt;li style="margin-bottom: 12px;"&gt;&lt;strong&gt;Output correctness:&lt;/strong&gt; schema parity, row counts, and any business-logic assertions your existing test suite covers&lt;/li&gt; 
&lt;/ul&gt; 
&lt;p&gt;The parallel-run phase is instrumented via Prometheus and Grafana, giving you side-by-side utilization metrics: raw infrastructure telemetry from your own environment, not a vendor-provided summary.&lt;/p&gt; 
&lt;p&gt;The typical engagement runs two to eight weeks depending on your organization's complexity, with about two weeks of engineering time. You get empirical data quickly that incorporates Falcon's performance, license cost, infrastructure, maintenance, and support requirements so you can make a data-informed decision on what's right for you. If the numbers don't clear your bar, there's no commercial conversation to have.&lt;/p&gt; 
&lt;p&gt;The commercial data analytics customer described above achieved 93% less compute infrastructure and 85% TCO reduction on a multi-terabyte-per-day real-time streaming ETL workload. A second benchmark, covering multi-terabyte daily batch and stream processing replacing production Databricks jobs across two pipelines, produced an 80% faster runtime, 94% lower cloud cost, and 63% lower TCO versus Databricks, implemented in 7 days. The first production customer recovered their full annual license cost from TCO savings within 73 days of cutover.&lt;/p&gt; 
&lt;p&gt;Once the benchmark clears your bar, cutting a live pipeline over without downtime is its own sequence. We've broken that out into &lt;a href="https://blog.haevek.com/databricks-offload-playbook" style="color: #f26b1d;"&gt;The Databricks Offload Playbook&lt;/a&gt;.&lt;/p&gt; 
&lt;h2&gt;Reading the Numbers: What Falcon Does to Your Annual Bill&lt;/h2&gt; 
&lt;p&gt;&lt;img src="https://blog.haevek.com/hs-fs/hubfs/Haevek-Falcon-PieChart.png?width=2544&amp;amp;height=1518&amp;amp;name=Haevek-Falcon-PieChart.png" width="2544" height="1518" alt="Haevek-Falcon-PieChart" style="height: auto; max-width: 100%; width: 2544px;"&gt;Consider a representative mid-market account with a $400K annual Databricks bill. If 75% of that spend is production compute — batch ETL, streaming pipelines, scheduled inference — that is $300K running on the JVM that can be migrated. Based on the 85% TCO reduction from the commercial streaming benchmark (which accounts for Falcon's license cost), 85% of $300K translates to roughly $255K in annualized savings. The remaining $100K covering interactive exploration stays on Databricks and is untouched.&lt;/p&gt; 
&lt;p&gt;Teams running 5–10x excess capacity see the largest absolute savings because the efficiency gain applies to the biggest base. Falcon's licensing is priced by parallel cores and throughput tiers, which means the cost model scales down with the infrastructure footprint. You are not paying for capacity you are no longer consuming.&lt;/p&gt; 
&lt;p&gt;Those savings right now are compelling, but it also sets an organization on a powerful path for the future. The expectation that enterprises will utilize their data for more AI inference will mean more utilization of Falcon instead of Databricks, where the overprovisioning problem would compound. Spark wasn't built for matrix operations at scale, and JVM GC behavior is especially punishing for memory-intensive workloads. Falcon handles batch ETL, streaming, and ONNX inference in a single pipeline, eliminating the separate model-serving infrastructure that many teams run alongside Databricks as inference volumes increase. Each new inference pipeline that spins up on the JVM is another candidate for the shortlist.&lt;/p&gt; 
&lt;h2&gt;A Next Step&lt;/h2&gt; 
&lt;p&gt;If you are an operator for your Databricks instance, you can find out in five minutes what the potential is for savings. Open a Databricks SQL notebook and run the &lt;code style="background: #F2F3F5; padding: 2px 6px; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.9em;"&gt;system.billing.usage&lt;/code&gt; query from the Prioritizing Workloads section above. Filter to the past 30 days, scope to JOB and STREAMING SKUs, and sort descending by total DBUs. Pull the top three results: job name, SKU, and total DBU consumption.&lt;/p&gt; 
&lt;p style="font-size: 1.15em; font-weight: 600; color: #0b0c0e;"&gt;That list is your Test Flight candidate set.&lt;/p&gt; 
&lt;p&gt;The Falcon Test Flight runs up to three of your production workloads side-by-side against your live Databricks jobs, in your environment, on your actual data, using your infrastructure metrics. Success criteria are defined upfront and tailored to your specific workloads. Within weeks, you'll have empirical numbers you generated yourself, not a vendor benchmark, that tell you exactly what Falcon would do for your stack.&lt;/p&gt; 
&lt;div style="background: #0B0C0E; padding: 32px; margin: 32px 0; text-align: center;"&gt; 
 &lt;p style="margin: 0 0 12px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Get started&lt;/p&gt; 
 &lt;p style="margin: 0; font-size: 1.2em; line-height: 1.5; color: #ffffff;"&gt;&lt;a href="https://haevek.com/contact" style="color: #f26b1d; font-weight: 600;"&gt;Contact Haevek&lt;/a&gt; to learn more or sign up for a test flight.&lt;/p&gt; 
&lt;/div&gt; 
&lt;div style="border-top: 3px solid #F26B1D; padding-top: 24px; margin: 40px 0 0 0;"&gt; 
 &lt;p style="margin: 0 0 16px 0; font-family: ui-monospace,SFMono-Regular,Menlo,monospace; font-size: 0.75em; letter-spacing: 0.12em; text-transform: uppercase; color: #f26b1d;"&gt;Keep reading&lt;/p&gt; 
 &lt;p style="margin: 0 0 12px 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/databricks-offload-playbook" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;The Databricks Offload Playbook: Four Steps to Cut Over Without Downtime&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;The operational sequence for moving a live pipeline off Databricks compute without an outage.&lt;/span&gt;&lt;/p&gt; 
 &lt;p style="margin: 0; line-height: 1.6;"&gt;&lt;a href="https://blog.haevek.com/databricks-falcon-faq" style="color: #f26b1d; font-weight: 600; text-decoration: underline;"&gt;Replacing Databricks Compute with Falcon: Frequently Asked Questions&lt;/a&gt;&lt;br&gt;&lt;span style="color: #666666;"&gt;Unity Catalog, Photon, Kubernetes support, benchmark access scope, and what stays on Databricks.&lt;/span&gt;&lt;/p&gt; 
&lt;/div&gt;  
&lt;img src="https://track.hubspot.com/__ptq.gif?a=48500980&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fblog.haevek.com%2Fposts%2Fslash-databricks-compute-bill-without-migrating&amp;amp;bu=https%253A%252F%252Fblog.haevek.com%252Fposts&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Apache Spark</category>
      <category>Cost Optimization</category>
      <category>Falcon Compute Platform</category>
      <category>Databricks</category>
      <category>Data Engineering</category>
      <pubDate>Wed, 12 Aug 2026 17:06:50 GMT</pubDate>
      <author>info@haevek.com (Haevek)</author>
      <guid>https://blog.haevek.com/posts/slash-databricks-compute-bill-without-migrating</guid>
      <dc:date>2026-08-12T17:06:50Z</dc:date>
    </item>
  </channel>
</rss>
