LTV Has a Blind Spot

We stress-tested LTV against 35,512 real customer cohorts. The prediction errors followed clear, repeatable patterns.

Executive Summary

LTV is arguably the single most important SaaS metric because it predicts the future: the total revenue a customer or cohort will generate. But does reality match its prediction? If not, it is causing faulty strategic decisions at thousands of companies right now.

LTV is a simple formula: ARPA divided by churn rate. Enter two numbers and get one number that estimates the total value of a customer relationship. It's one of the foundational metrics in SaaS, shaping acquisition budgets, payback expectations, fundraising conversations, and product strategy.

But when you compare the prediction to reality, how accurate is it?

To answer that question, we stress-tested LTV against reality across 35,512 customer cohorts from 3,331 ChartMogul company accounts. Specifically, we asked: if you calculated LTV when a cohort formed, how closely did that prediction match reality 12, 24, and 36 months later?

The headline is counterintuitive: LTV prediction errors are not random. They are systematic, making the metric reliable in some situations and consistently misleading in others.

Key findings

  • Prediction errors are systematic, not random. E-Commerce businesses consistently overestimate customer value, while enterprise-heavy ones often underestimate it because customer size, retention, and expansion interact differently.
  • Many cohorts are substantially wrong. 28.3% of cohorts differ from reality by more than 50%.
  • The median error is small only because opposing errors often cancel each other out. LTV is frequently too high and too low by similar amounts, making the overall median appear accurate.
  • Most prediction error comes from one assumption: using company-wide ARPA to estimate new customer revenue.
  • Lower-paying customers counterintuitively often stay longer than higher-paying customers, but expansion from the survivors eventually flips LTV from overestimating to underestimating revenue.
  • LTV is best used as a directional metric, not a precise forecast. Cohort revenue curves and net MRR movements provide the missing context.

Instead of asking whether LTV is "right," we asked where it goes wrong. What causes the errors? How do they compound? How do they change over time? And what do they mean for how you should use LTV?

Written by Thomas Anastaselos, with research guidance and editorial feedback from Jason Cohen.

Thomas Anastaselos

Thomas Anastaselos is Director of Revenue Operations & Data Analytics at ChartMogul. He leads initiatives across revenue operations, go-to-market strategy, product-led growth, and product analytics, helping turn complex data into practical insights that drive better business decisions.

Jason Cohen

Jason Cohen is the founder of WP Engine, one of the world's largest WordPress hosting companies, and the author of Hidden Multipliers. He writes about SaaS metrics, growth, and pricing at asmartbear.com.

What the LTV formula assumes

LTV = ARPA ÷ Churn Rate is a two-variable model. Its simplicity comes with constraints.

To understand where it goes wrong, start with what the formula assumes:

Every new customer pays the same as the company average. ARPA is a blended number. It represents every subscriber on your account right now, small and large, new and long-tenured, expanding and contracting. When you use it to predict the revenue a new cohort will generate, you're assuming the customers who just signed up look like your entire current customer base, including customers who have been with you for three years and expanded their spend multiple times.

The churn rate stays constant. The trailing six-month average customer logo churn rate is applied uniformly to every new cohort regardless of who they are, how they were acquired, or how well they fit the product, or the fact that some customers already churned.

Revenue per customer stays flat. Expansion or contraction is not taken into account specifically, just indirectly via the blended ARPA. Whatever customers pay when they arrive, they pay forever, and customers who canceled early pay the same on average as customers who have been with the company for five years.

These are known trade-offs. Simpler models are easier to understand and harder to game. The question is whether those trade-offs matter in practice. Let's look at the data.

How we tested it

For each ChartMogul account in the dataset, we defined cohorts: all customers who signed up in a given calendar quarter, identified by when they first became paying customers. We analyzed 20 quarterly cohorts, Q1 2020 through Q4 2024, across all 3,331 accounts, which resulted in 35,512 (account, cohort quarter) pairs in total.

At the moment each cohort was created, we recorded the account's ARPA and trailing six-month churn rate, exactly as the LTV formula does. We then computed the revenue predicted by the LTV formula for those customers over 12, 24, and 36 months.

Next, we measured what those same customers actually generated. For each customer, we tracked their MRR month by month. We summed actual MRR over each time window. We excluded cohorts whose observation window extended past the data cut-off date (June 2026) to avoid censoring bias.

This gave us a direct comparison between predicted and actual revenue for the same customers over the same time period, using only the information that was available when they first became paying customers.

The formula looks quite accurate

The typical cohort generates 2.4% less revenue over 12 months than LTV predicts. That sounds reassuring.

However, 28.3% of all cohorts have a prediction error larger than 50% in either direction.

A box plot titled 'Distribution of LTV Prediction Errors at 12 Months'. The box spans from the 25th percentile at -28% to the 75th percentile at +28%, with the median at -2.4%. Shaded zones on either end mark the 28.3% of cohorts with prediction errors beyond plus or minus 50%.

The formula uses a blended ARPA that includes years of expansion by existing customers. Those customers have been adding seats, upgrading plans, and increasing their spend for months or years. New customers haven't done any of that yet. The formula assumes they already look like the average customer, even though they haven't had time to become one.

Two inputs = two ways to be wrong

The formula has two inputs, so it can be wrong in two places: the ARPA in the numerator, or the churn rate in the denominator. Both are wrong, but in different directions, by different amounts, and for different reasons.

Failure mode 1: New customers aren't average customers.

The formula uses the company's blended average revenue per customer. New customers nearly always pay less than that average. They've just signed up, which means no seat expansions, no plan upgrades, no multi-year growth. The blended average includes all of that. So the formula assumes new customers generate revenue they haven't earned yet.

At the median, this ARPA mismatch has a modest effect on prediction accuracy: -4.9%. But the median hides how much the mismatch varies across cohorts. When we decomposed each cohort's prediction error into ARPA and retention components, ARPA mismatch explained 84.9% of the variation in prediction accuracy. Put differently: the typical ARPA effect is small, but differences in ARPA mismatch are one of the main reasons LTV predictions are more accurate for some cohorts than others.

Failure mode 2: Cohorts don't churn at the average rate.

The formula uses the company's trailing six-month churn rate to project how many customers will survive each month. But a specific cohort can churn at a very different rate. New enterprise buyers, SMB converters, and trial-to-paid users don't share the same survival curves, even if they all joined in the same quarter. The company average tells you little about the specific cohort in front of you.

At the median, these two errors happen to partially offset each other (ARPA and churn). That is why -2.4% looks small.

The survival paradox

One of the most surprising findings is which customers actually stay longest.

Conventional SaaS wisdom says higher-paying customers stay longer. Enterprise customers have lock-in, multi-year contracts, and dedicated account management. Cheaper customers have more alternatives, lower switching costs, and less organizational gravity. At WP Engine, for instance, this pattern applies: larger companies are stickier.

ChartMogul's dataset skews toward SMB and includes a substantial B2C segment. Within that dataset, the pattern reverses. We split customers into four groups based on their signup MRR and measured how many were still active after 12 months. The result was surprising: 12-month retention rates fall as relative price rises. The cheapest quartile survived at 74%; the most expensive survived at only 45%.

A second check told the same story: the median signup MRR of still-active customers falls over time, not rises. Within a cohort, the higher-paying customers tend to leave first. This may partly reflect the "become what you hunt" pattern described in our previous research: many SaaS companies successfully acquire SMB customers before pushing upmarket, where retaining larger customers becomes more challenging.

A bar chart titled '12-Month Survival Rate by Signup MRR Quartile'. Survival rate falls as signup MRR rises: the cheapest quartile of customers survives at 74% after 12 months, while the most expensive quartile survives at only 45%.

Expansion, not survival, inflates ARPA

At first glance, this seems to contradict the ARPA story. If cheaper customers survive longer, shouldn't cohort ARPA fall over time rather than rise? One possible explanation is that in SMB-heavy markets, cheaper subscribers have little reason to leave. The software is embedded in their workflow, the cost is below the threshold of a deliberate cancellation decision.

Higher-paying companies are more deliberate about ROI. When the value isn't clear, they switch. The enterprise lock-in dynamic that makes larger B2B customers sticky doesn't apply to the $500/month SMB buyer or the $29/month consumer.

What this means for ARPA. If cheaper customers survive longer, then as a cohort ages, its average MRR should fall. That's the opposite of ARPA inflation, and it appears to be what happens at the customer level.

So where does company-level ARPA inflation come from? Expansion. Customers who stay put upgrade. We measured this directly: 45.6% of current MRR from customers who survived from 2020-2021 comes from expansion after signup, not from their original subscriptions. The blended ARPA is inflated because survivors expand.

What does this mean for LTV? The ARPA used to predict new cohort revenue is inflated by expansion your existing customers have already experienced, but new customers haven't had time to achieve yet.

Up to this point we've explained why LTV initially overestimates new customer value. But those sources of error aren't constant. As cohorts mature, customers begin expanding. That gradually makes the initial ARPA mismatch less important while expansion becomes more important. You can see the shift when we compare the same cohorts over longer time horizons.

How the prediction changes over time

Early on, blended ARPA makes LTV too optimistic. Over time, customer expansion overtakes that initial bias. At 24 months, the median prediction error is +6.9% and at 36 months it's +14.4%. By 24 months, the LTV is understating revenue instead of overstating it.

The reason is expansion. At 12 months, ARPA mismatch and expansion roughly cancel each other. By 24 months, enough customers have upgraded seats and plans that expansion wins. By 36 months, cumulative expansion pushes the number even further.

The spread of outcomes also nearly doubles by 36 months, making LTV much less reliable.

A chart titled 'LTV Prediction Error Over Time'. The median prediction error moves from -2.4% at 12 months to +6.9% at 24 months and +14.4% at 36 months, while the P25-P75 spread of outcomes nearly doubles by 36 months.

Not every business behaves this way

The overall pattern, i.e. cheaper customers stay longer, is real but not universal. Not every company behaves this way, so let's look where it does and where it doesn't.

For each company, we measured the correlation between customer signup MRR and 12-month survival. Positive correlation means bigger customers stay longer, negative means bigger customers churn sooner.

Here's what we found:

Group Share of companies Pattern
Negative 37.2% Bigger customers churn sooner
Neutral 48.9% No meaningful relationship
Positive 13.9% Bigger customers stay longer

The expected enterprise pattern, i.e. bigger stays longer, holds for only 14% of companies. For 37%, the inverse is true.

A chart titled 'ARPA-Retention Correlation Groups'. 37.2% of companies show a negative correlation where bigger customers churn sooner, 48.9% show no meaningful relationship, and 13.9% show a positive correlation where bigger customers stay longer. B2C companies make up 37.8% of the negative group but only 6.6% of the positive group.

B2C companies make up 37.8% of the negative group, but only 6.6% of the positive group. This fits the intuition for many consumer businesses: the product is embedded in daily life, and the cost is too low to justify cancelling. Higher-paying subscribers on the other hand are more deliberate and more likely to churn when value isn't clear. Small B2B companies follow the same pattern: 77% of the negative-correlation group has under $500k in ARR.

The positive group (bigger customers staying longer) is 93% B2B and dominated by larger companies. This is the classic enterprise pattern: higher MRR comes with greater integration, higher switching cost, and deeper organizational dependency.

These differences in customer behavior show up directly in differences in LTV accuracy.

Group Median LTV error % cohorts unreliable
Negative - bigger churn sooner +1.3% 15%
Neutral -1.3% 25%
Positive - bigger stays longer -5.7% 45%
A chart titled 'LTV Accuracy by ARPA-Retention Correlation Group'. The negative-correlation group, where bigger customers churn sooner, has the lowest median error at +1.3% and the fewest unreliable cohorts at 15%. The positive-correlation group, where bigger customers stay longer, has the highest median error at -5.7% and the most unreliable cohorts at 45%.

LTV is most accurate when bigger customers churn sooner, and least accurate when they stay longer.

The model is most accurate when its two mistakes offset each other: it overstates the value of new customers, but those higher-paying customers also churn faster, bringing actual revenue closer to the prediction. It is least accurate when the mistakes compound: higher-paying customers start out more valuable and also stay longer, so the formula underestimates them twice: first when they sign up, and again as they expand over time.

Who it fails most

The median error isn't evenly distributed. The biggest differences appear across industry, company size, and business model.

By industry:

A bar chart titled 'Median LTV Prediction Error by Industry'. Bars show the median 12-month prediction error by industry, from -10.9% for E-Commerce to +0.7% for Education, with right-side labels showing the share of cohorts wrong by more than 50%, ranging from 22.8% to 31.3%.
Industry Median 12m error % cohorts wrong by >50% ARPA ratio Retention factor
E-Commerce -10.9% 29.0% 0.88 1.02
Workplace & Productivity -6.7% 29.8% 0.90 1.06
Infra / Dev Tools -5.9% 31.3% 0.91 1.06
Business Intelligence -3.2% 24.5% 0.87 1.18
Sales, Marketing & CS -3.1% 28.4% 0.91 1.10
Finance & Operations -2.2% 30.5% 0.91 1.10
Membership -1.7% 22.8% 0.99 1.01
Education +0.7% 23.1% 1.02 0.99

E-Commerce has the largest overprediction. New customers pay well below the average and show little expansion, leaving little to offset the initial ARPA mismatch.

At the other end of the table we see the opposite story. Business Intelligence shows the largest expansion effect (1.18x) as analytics tools become more embedded over time, and customers upgrade as usage grows. Industries with strong post-signup expansion consistently show smaller LTV errors despite similar ARPA mismatches.

The differences appear to reflect business models more than industries. LTV is most accurate when new customers look similar to the existing customer base and expand after signup. It's least accurate when new customers start well below the average and never close the gap.

By company size:

MRR band Median error % unreliable ARPA ratio Retention factor
Over $30M ARR -19.1% 19.9% 0.69 1.21
$10M-$30M ARR -6.8% 25.1% 0.83 1.11
$500k-$10M ARR -4.3% 27.5% 0.90 1.09
Under $500k ARR +0.5% 29.5% 1.00 1.03

The largest companies ($30M+ ARR) show the biggest overprediction but the fewest large misses. Their errors are large at a -19% median, but consistent. Companies have been around long enough for established customers to expand dramatically, inflating the blended ARPA well above what new cohorts actually pay. But those new customers also tend to expand substantially after signup, partially closing the gap over time.

Small companies (under $500k ARR) now show a slight underprediction (+0.5% median error) as the formula slightly underestimates revenue for the smallest companies, possibly because their expansion trajectories outpace what the average implies. But they also have the most unreliable cohorts (29.5%). Their errors are noisy rather than systematic. On average, the formula is about right, but individual cohort predictions vary widely.

By business model:

Target market Median error % unreliable ARPA ratio Retention
B2B -5.5% 30.2% 0.88 1.10
B2C +0.4% 20.2% 1.03 0.96

B2B and B2C present fundamentally different accuracy profiles.

For B2B businesses, the biggest issue is ARPA mismatch. Existing customers have expanded, pushing the average above what new customers pay.

B2C companies are essentially neutral in aggregate (+0.4% median error) and have the lowest rate of badly-wrong cohorts (20.2%). B2C pricing is simpler and more uniform, which means the formula's core assumptions hold better. The formula's blind spot is smaller in consumer markets.

You can't just add a warning label

We tested this using the strongest signal available when a cohort is created: the ARPA ratio. It measures how different a new cohort's revenue per customer is from the company-wide average used in the formula. Because it is observable immediately, it can be used before any retention data exists.

The logic is simple. If new customers are paying substantially more or less than the average, the formula's ARPA assumption is already wrong. That alone can make the resulting LTV estimate unreliable.

We tested different thresholds for flagging these cohorts and measured the trade-off between precision and recall: how often the flag correctly identified a bad estimate, and how many bad estimates it caught.

Flag threshold Precision Recall Companies flagged
0.25 39.7% 66.7% 47.6%
0.35 (best balance) 45.3% 57.6% 36.7%
0.50 54.8% 45.5% 23.6%
1.00 81.0% 9.6% 3.3%

To catch 57% of badly-wrong predictions at the best precision-recall balance, you would flag 37% of all companies. Raising the threshold to catch only the most extreme cases (threshold 1.00) sharpens precision to 81%, but only catches 10% of bad predictions.

Which reflects the underlying reality: new customers are rarely identical to the average, so the ARPA assumption is wrong to some degree for almost every company. A warning label would end up flagging a large share of companies, making it neither actionable nor especially credible.

What this means for how you use LTV

There are three ways to respond to what the data shows: adjust how you read your current number, fix the formula itself, or lean on complementary metrics.

Adjust how you read your current number.

Treat LTV as directional, not precise. The median error is small, but the distribution around it is wide. For any specific cohort, the actual outcome could easily land 25-30% above or below the prediction. If your LTV:CAC target already includes a safety margin, remember some of that margin may already be absorbed by normal LTV uncertainty.

A few segment-specific adjustments, based on what the data shows:

  • E-Commerce: the ARPA assumption is systematically too optimistic here (median error -10.9%). Consider discounting LTV by 10-15% before using it for acquisition decisions.
  • Large companies ($30M+ ARR): the error is predictable and one-directional. LTV tends to overstate by roughly 20%, largely because blended ARPA bakes in years of existing-customer expansion.
  • Enterprise-heavy quarters: if new customers are materially larger than your existing base, expect the opposite problem. LTV will likely understate what they're worth.

One caveat: these describe the typical company in each segment, not every company. Before applying any adjustment, check how closely your own historical LTV predictions matched reality.

Fix the formula itself.

The highest-impact structural fix is replacing company-level blended ARPA with the actual signup-time revenue of the new cohort. Since ARPA mismatch drives the largest share of error variance, this alone would meaningfully improve accuracy, and the data needed (MRR at signup) already exists.

Expansion is harder to fix this way. It isn't knowable at signup, and it accounts for a large share of the remaining error, so a corrected formula would still miss the second, larger source of drift. There's also a practical cost: a cohort's realized ARPA isn't finalized until the cohort closes, which would delay the calculation and add complexity.

Use complementary metrics.

For near-term unit economics, payback period may be a more reliable substitute. It's bounded (typically 12-24 months rather than an infinite horizon) and sensitive to the variables that actually move (margin, early churn, CAC), though a single expensive acquisition push can temporarily distort it.

For a longer horizon, cohort revenue curves show what past cohorts actually generated, not a projection, but a historical record. A rising 24-36 month slope signals healthy retention and expansion; an early flattening or decline is a visible, dateable problem.

Net MRR movements complement both by explaining why those curves look the way they do, breaking out expansion, contraction, churn, and reactivation, where LTV compresses all of it into one number.

No single view is complete. Together, they give a far more accurate picture of customer value than LTV alone.

The closing question

LTV is limited because it compresses two unstable assumptions into a single number. Understanding those assumptions matters more than trusting the number itself.

The formula will remain useful, as it is simple enough to explain in a board meeting, easy to track over time, and good at summarizing a business in a single number. Used directionally and alongside other metrics, LTV remains a valuable tool.

Methodology & data

Methodology notes

  • Dataset: ChartMogul platform data from 3,331 company accounts, covering Q1 2020 through Q4 2024 and producing 35,512 (account, cohort quarter) pairs. Cohorts are defined as all customers whose first new-business movement occurred within the same calendar quarter.
  • Prediction: When each cohort formed, we calculated LTV using ChartMogul's formula based on the account's blended ARPA and trailing six-month logo churn rate. We then computed the LTV-implied revenue prediction for the following 12, 24, and 36 months.
  • Actuals: For each customer, we tracked MRR month by month and summed actual MRR over each time window. Cohorts whose observation window extended past the data cut-off (June 2026) were excluded to avoid censoring bias.
  • Error metric: Prediction error is (actual - predicted) / predicted × 100. Negative means the formula overpredicted and positive means it underpredicted. Unreliable cohorts are those with absolute error greater than 50%.
  • Variance decomposition: We split each cohort's prediction error into two components: ARPA mismatch and retention mismatch. We then measured how much of the variation in prediction accuracy across cohorts each component explained. The two components, plus their interaction, account for 100% of the explained error.
  • ARPA-retention correlation: For each account, we measured the relationship between customer signup MRR and 12-month survival. Accounts where larger customers tended to stay longer were classified as positive; those where they churned sooner were classified as negative; the remainder were classified as neutral. A minimum correlation threshold was applied to reduce noise from small samples.

A note on data

This article reflects an analysis conducted on ChartMogul's live platform data as of June 2026. Earlier versions of this work showed slightly different median errors and variances, but the fundamental points and factors that affect them remain the same. The changes are due to underlying datasets evolving and company accounts added or removed from the dataset.

Subscribe to The SaaS Roundup

Join 24,000 others who get handpicked industry reads and the latest insights directly to their inbox every Friday.