The Consumption Gap: How AI Is Splitting SaaS Pricing Into Two Layers

by | Jun 12, 2026 | Business, SaaS Growth Hacks

Snowflake closed fiscal 2025 with $3.6 billion in product revenue, and roughly 95 percent of it was metered, not licensed. Datadog, on a different consumption clock, posted about 120 percent net revenue retention on $3.43 billion in 2025 sales. Neither sells a seat. Both grow when their customers do more work inside the product, not when procurement re-counts headcount.

Now look at what AI is doing to the rest of SaaS. 78 percent of IT leaders saw an unexpected consumption charge in the past year, and roughly 63 percent of enterprises blew through their AI budget by 30 percent or more in year one. 38 percent of SaaS companies now run usage-based pricing, up from 27 percent in 2021, with 61 percent on hybrid models that pair a subscription floor with a metered upside.

That is the consumption gap. The price the customer pays per month is one number. The work the software does on the customer’s behalf is another, and AI has pushed the two further apart than at any point since the seat replaced the perpetual license. The SaaS companies winning in 2026 are rebuilding the price card in two layers: a stable subscription that anchors the relationship, and a consumption layer that catches value when the product actually runs.

What the consumption gap actually is

Every SaaS pricing model lives inside a value metric. Salesforce picked the seat. AWS picked the compute-hour. Twilio picked the message. The metric works when it correlates with delivered value, and the seat held up for two decades because serving one more user of a multi-tenant app was close to free.

AI broke that math. As a16z partners put it on a recent podcast, the unit of value in AI software is no longer the user, it is the output. A code assistant generates suggestions. A support bot resolves tickets. An agent runs workflows. Each output has a real variable cost behind it, paid in tokens, GPU minutes, and tool calls.

The mismatch is structural. The customer is invoiced in a flat number per seat, but the back-end is metered in tokens. When usage spikes, the vendor eats the margin. When usage drops, the customer pays for capacity they never touched. Bessemer’s AI pricing playbook tracks the moment AI-native companies started abandoning seats in favor of usage, output, and outcome models that move with the actual work.

The two-layer stack: subscription floor, consumption upside

The clearest example sits inside Intercom. A seat license starts at $29 per agent per month, and Fin, its AI agent, layers on top at $0.99 per resolution. One layer pays for the inbox, the routing, the reporting, the user count. The other pays for the work Fin does when it closes a conversation without a human.

The pattern is repeating across the AI SaaS stack. Chargebee’s 2025 State of Subscriptions found 43 percent of companies on a hybrid model, and OpenView expects 61 percent by year-end. Vendors are not abandoning subscription. They are putting a meter behind it.

There is a financial reason and a buyer reason for the split. Pure consumption revenue moves with customer activity, which is great for expansion and miserable for forecasting. The CFO needs a base that looks like ARR for the audit, the covenants, and the board. The buyer reason is just as concrete: Bessemer notes that consumption-only pricing works for technical buyers who already understand the unit. For everyone else, it is a translation problem.

Layer one: the subscription floor

The base subscription does three jobs. It funds the platform and the workflow surface. It gives the buyer a predictable number procurement can sign. And it produces revenue that still walks and talks like ARR, which matters because public consumption businesses still get measured on remaining performance obligations and committed contract value, not just trailing consumption.

Operator caveat: in seat-based models above roughly 500 seats, the floor needs an admin and security tier, otherwise the customer will push the per-seat number down at renewal. The floor cannot be a single SKU at scale.

Layer two: the consumption upside

The consumption layer is where AI work gets paid for. Outputs, agent resolutions, API calls, tokens, document drafts. The vendor picks the unit that correlates with delivered value, then meters it. Snowflake metering compute and storage independently is the canonical version, explained in their architecture writeups, and it produced the 158 percent NRR Snowflake has posted at points in its history. Datadog metering hosts, custom metrics, and logs is the same idea applied to observability. The unit changes, the principle does not.

Why billing infrastructure became a strategic line item

A two-layer model is easy on a whiteboard and brutal in production. The vendor has to meter every token, every API call, every agent step, apply tiered pricing without latency, handle prepaid credits, refunds, and tax across geographies, and produce an invoice the customer’s finance team will not bounce.

Stripe rebuilt Billing around exactly this problem and in March 2026 launched a private preview for LLM token billing. Stripe auto-syncs per-token prices for OpenAI, Anthropic, and Google models, lets the vendor add a percentage margin, and emits meter events behind a single API. PYMNTS framed the move as Stripe arguing that subscription needs a usage-based upgrade, not a replacement.

Behind Stripe, an entire metering category has emerged. Paddle, Metronome, Orb, and m3ter sell variations of the same thing: a high-volume meter, a rules engine, and an invoice generator that can survive an audit. They got funded because building this in-house takes 3 to 6 months of engineering work and then never stops needing maintenance.

Billing has stopped being a back-office function. It is a wedge. A SaaS company that can ship a new SKU, a new meter, or a new credit bundle in days instead of quarters can price into demand faster than its competitors. That speed compounds into NRR.

Why the NRR math favors the two-layer stack

Net revenue retention is where the gap shows up on the income statement. Companies with usage-based or hybrid models report 120-plus percent NRR against roughly 110 percent for subscription-only peers, and median logo churn drops from 12 percent to 8 percent in the same dataset.

The mechanism is structural. In a flat subscription, every dollar of net new revenue from an existing customer needs an upsell motion. In a metered model, the dollar arrives the moment the customer does more work. SaaS Mag framed NRR as the defining SaaS metric of 2026 earlier this year precisely because consumption changed what organic expansion can look like.

The risk is symmetrical. If a customer’s usage drops, the metered revenue drops with it, and the vendor has no contractual claim on the lost dollars. That is why the subscription floor matters. It sets a minimum the vendor can plan against, and gives the renewal conversation something to anchor on when usage swings.

Three SaaS categories, three two-layer designs

Customer support: Intercom Fin

Intercom’s pricing card now reads in two columns. A seat license for the inbox and routing, and a per-outcome charge for Fin, which only fires when the AI agent resolves the conversation. Reported resolution rates run 42 to 50 percent in published case studies, which means roughly half of inbound tickets generate a consumption charge. The base pays for the platform, the consumption layer pays for the work.

Data and observability: Snowflake and Datadog

Snowflake’s product revenue is roughly 95 percent consumption, with credits as the unit. Datadog charges per monitored host, per million custom metrics, per billion logs ingested. Both still offer committed contracts that lock customers into a minimum spend, the floor under the meter. Battery Ventures’ OpenCloud research has documented how the committed-spend floor is what lets these companies trade as ARR businesses, not as commodity utilities.

Productivity and AI copilots

Productivity SaaS is the messiest category, because the seat is hard to retire without breaking buyer expectations. The emerging pattern is a seat for the editor or workspace plus AI credits for the copilot. Tomasz Tunguz has noted that the right metric in production is intelligence per dollar, value per token spent. Credit bundles try to expose that without forcing the customer to count tokens.

What changes for the SaaS finance team

Forecasts go bottoms-up by meter, not top-down by seat. The CFO models token consumption, agent run-rate, and feature adoption on top of headcount. The better SaaS finance teams in 2026 run monthly meter forecasts the same way they used to run pipeline, with confidence bands and a designated owner per meter.

Deferred revenue stops being a clean liability. Prepaid credits can be drawn down in any pattern, which changes revenue recognition, the cash-to-revenue gap, and how auditors approach the close. Hybrid finance teams are rethinking what healthy growth looks like when recurring revenue fluctuates with consumption.

Gross margin moves with the product mix. Every dollar of AI consumption carries an inference cost behind it, so company-wide gross margin now depends on how much revenue comes from the consumption layer. SaaS Mag covered the AI COGS problem in detail in May, and the consumption gap is the demand-side mirror of the same picture.

How to draw your two layers without breaking the sales motion

A few rules have emerged from the SaaS companies that shipped a two-layer model in the last 18 months.

Anchor the floor in something the buyer already understands. Seats, workspaces, environments, named users. The floor’s job is to be legible to procurement, not optimal for unit economics. Pick the unit your buyer already pays for elsewhere.

Meter the value, not the cost. If the unit is tokens, the customer is forced to think about what the vendor pays, not what the customer gets. Wherever possible, meter a unit the customer connects to value: resolutions, leads enriched, documents drafted, deployments shipped. Token-level meters belong inside the dashboard, not on the invoice.

Bundle credits to absorb shock. Prepaid credits give the customer cost predictability and the vendor better cash conversion. Credit models are now used by 79 of the top 500 SaaS companies, up 126 percent from 2024.

Engineer overage with care. The single biggest source of churn around consumption pricing is invoice shock. Soft caps, alerts at 75 and 90 percent of plan, and an automatic note when usage trends suggest an overage are now table stakes. Done well, overage becomes an expansion conversation instead of a churn event.

Operator caveat: the in-app usage dashboard is part of the pricing. Vendors that ship a real-time usage view retain better than vendors that ship an itemized PDF after the fact.

The opportunity hiding inside the gap

The same shift handed SaaS founders something the per-seat model could never deliver: automatic expansion. The dollar arrives the moment the customer does more work. No upsell meeting, no QBR, no rebadging exercise. Battery Ventures put it bluntly: consumption is the next evolution of subscription, not a replacement, and it produces a more natural land-and-expand motion than seats ever did. The seat is not dead. It is the floor. The new ceiling is whatever the product can do.

FAQ

What is the consumption gap in SaaS pricing?

The consumption gap is the structural difference between what the customer is billed and what the software actually does on the customer’s behalf, especially in AI-heavy SaaS. A per-seat subscription is flat. AI workloads, agent steps, and tokens are variable, so the gap widens when usage spikes outrun what the seat price was built to cover. The fix most SaaS companies have settled on in 2026 is a two-layer price card: a subscription floor that funds the platform and a consumption layer that scales with the work the product performs.

How is consumption-based pricing different from usage-based pricing?

The terms are used interchangeably in most operator writing, and OpenView treats them as one category. The lighter distinction: usage-based pricing meters a defined unit such as API calls or active seats, while consumption-based pricing meters resources drawn down, like compute hours, tokens, or credits spent. Both describe a model where the customer is billed for what they used in a period, often on top of a base subscription, rather than for the right to use the product.

What NRR can a SaaS company expect from a hybrid pricing model?

OpenView’s annual benchmarks put usage-based and hybrid SaaS companies at roughly 120 percent median net revenue retention, against about 110 percent for subscription-only peers. Public reference points run higher: Datadog at roughly 120 percent on multibillion-dollar revenue, Snowflake historically at 158 percent. The gain comes from expansion happening automatically as customers do more work in-product, instead of being negotiated at renewal. The risk is symmetrical: when usage drops, so does revenue, which is why a subscription floor matters.

Should an AI SaaS company price per token, per outcome, or per seat?

All three, usually layered. The seat or workspace works as a stable floor procurement can sign. The outcome unit, such as a resolved ticket or generated document, is the cleanest expression of value for non-technical buyers. The token meter belongs inside the dashboard, not on the invoice, because customers cannot estimate need in tokens. Intercom Fin’s $0.99 per resolution and Snowflake’s credit model are the two patterns most often copied in 2026.

Do investors value consumption SaaS differently from subscription SaaS?

Public-market data through 2026 suggests consumption-driven SaaS companies trade at higher revenue multiples on average, because expansion is built into the revenue motion. The catch is that pure consumption revenue is less predictable than ARR, so committed contracts and minimum-spend agreements matter for valuation. Battery Ventures’ research notes that the most valuable cloud companies still anchor their revenue on a recurring base. The two-layer model exists in part to keep that anchor while capturing consumption upside.

The bottom line

SaaS spent 20 years selling access. AI made the work itself ownable. The companies thriving in 2026 are the ones who priced both. A subscription floor for the access, a consumption layer for the work, and the billing infrastructure to make them feel like one product to the customer. The gap was never the problem. The two-layer answer is.

Hybrid and pure usage-based pricing are now the majority pattern in SaaS, per OpenView and Chargebee survey data, 2019 to 2026.
Hybrid and pure usage-based pricing are now the majority pattern in SaaS, per OpenView and Chargebee survey data, 2019 to 2026.
NRR widens once a company adds a consumption layer, with enterprise UBP/hybrid models reaching 120 to 160 percent.
NRR widens once a company adds a consumption layer, with enterprise UBP/hybrid models reaching 120 to 160 percent.
The deeper the AI integration, the larger the consumption layer needs to be to protect gross margin.
The deeper the AI integration, the larger the consumption layer needs to be to protect gross margin.

Recent Posts

Explore Topics