Home Blog Resume Contact Ask AI About Me
Home/Blog/When Klaviyo Predicted CLV Is Wrong About You…
ArticleKlaviyoKlaviyoPredicted CLVSegmentation

When Klaviyo Predicted CLV Is Wrong About Your Store

9 min readBy Miloš Mitrović

Predicted CLV quietly drives real money in a Klaviyo account: VIP tiers, win-back budgets, suppression rules and the seed audiences you push to paid. So a model that misreads your catalogue sends spend and attention to the wrong customers, and it does it silently. Klaviyo builds the model from your own order history and retrains it weekly, which serves high-repeat consumables well and considered, once-a-year catalogues badly. The number always renders. The question is whether it describes a buying pattern your customers actually have.

Key takeaways

  • Klaviyo shows predictive analytics only once at least 500 customers have placed an order, you hold 180 days of order history with orders in the last 30 days, and some customers have placed 3 or more orders.
  • Predicted CLV projects the next year of spend and, by Klaviyo's own statement, is reliable across a population rather than for any single profile.
  • The buy-till-you-die models that dominate non-contractual ecommerce infer future orders from a purchase rhythm, so they misread catalogues where a purchase is rare and considered.
  • A blank Predicted CLV on a profile means Klaviyo lacks the data to score that person, not that the person is low value.
  • On low-repeat catalogues, historic CLV, first-order value and product-level signals beat the predicted figure.

What does Klaviyo actually calculate when it shows Predicted CLV?

It projects the money a customer will spend over the coming year and derives that from their purchase frequency, average order value and time between orders. Klaviyo exposes three related figures, and the Klaviyo guide to segmenting on customer lifetime value defines them as Historic CLV, the value of past orders minus refunds, Predicted CLV, the projected spend over the next year, and Total CLV, the two added together. The model is not static. Klaviyo rebuilds the CLV model from your account data and retrains it at least once a week, so the figure moves as orders and engagement accumulate.

The load-bearing word is projected. Historic CLV is arithmetic you could reproduce in a spreadsheet. Predicted CLV is a forecast, and a forecast carries every assumption of the model that produced it. When operators treat the two as interchangeable, they inherit those assumptions without knowing they signed up for them.

When does Klaviyo have enough data to predict at all?

Only after your account clears three thresholds, and until then the predictive section stays blank instead of guessing. Klaviyo requires at least 500 customers who have placed an order, 180 days of order history with orders inside the last 30 days, and some customers who have placed 3 or more orders before predictive analytics renders at all. Cancelled, refunded and zero-value orders do not count toward the model, so a store padded with test orders or comped units clears the bar later than its dashboard suggests.

Two accounts can pass those thresholds and still get very different quality from the model. The 500-customer floor describes the population; whether a given person receives a score depends on that individual's own order count. Klaviyo states that a blank predictive block on a profile means it does not have enough data on that individual to make a prediction. Read a blank as unknown, never as low value, because suppressing or de-prioritising blank profiles throws away your newest and often most valuable buyers.

Why does Predicted CLV misfire on considered, low-repeat catalogues?

Because the models that dominate non-contractual ecommerce infer future purchases from a rhythm, and a considered catalogue has no rhythm to read. Klaviyo describes its method generically as machine learning retrained weekly and does not publish the exact algorithm, so treat any specific model name as inference. The wider field, though, is dominated by the buy-till-you-die family, and its assumptions explain the failure cleanly.

The Pareto/NBD specification documented by PyMC-Marketing assumes each active customer buys following a Poisson process, that purchase rates vary across customers on a gamma distribution, and that customers silently drop out at some unobserved point. Its common cousin, BG/NBD, treats a customer with no repeat purchase as still active until proven otherwise. Those assumptions fit a coffee, skincare or supplement brand, where a healthy customer reorders every few weeks and a long gap is a genuine warning.

They fall apart on a mattress, a sofa, an engagement ring or a premium appliance. A single large purchase followed by a year of quiet is ordinary behaviour for that catalogue, not decay, yet a frequency-driven model reads the silence as churn and marks the customer down. A second weakness compounds it. As the comparison of Pareto/NBD forecasting methods in the Journal of Business Economics and the broader literature note, these models forecast transaction counts and treat spend as a loosely coupled add-on, so a catalogue where one order can be fifty times another is exactly where the monetary estimate is weakest.

On a high-consideration catalogue the trustworthy signal is what a shopper browsed and weighed, not how often they buy. That is the reasoning behind rebuilding cart reminders around each shopper's considered products, and it applies just as directly to how you value them.

How can you tell the prediction is wrong before you act on it?

Compare the model's expectation against what your customers actually did, at the segment level, before you wire any prediction into a flow. Klaviyo publishes a churn risk score per profile, exported as a number between 0 and 1, and the Klaviyo churn model raises that score as elapsed time passes the customer's average buying cycle. On a low-repeat catalogue that logic pushes almost everyone toward high churn risk within months, which is your first tell: if most of your file lands in the top churn band, the model is describing your category, not your customers.

Three checks catch a misfit fast:

  • Build a segment of "expected date of next order in the past" and see what share of your file it captures. If it swallows the majority, the interval assumption does not hold for you.
  • Chart Predicted CLV against Historic CLV. On a well-fit catalogue the predicted figure adds a sensible forward slice; on a misfit it collapses toward zero for anyone past one buying cycle.
  • Hold out a cohort from 12 months ago, then compare last year's Predicted CLV against the spend that actually followed. A model that fits will track the population total within a reasonable band even while individual rows scatter.

The stakes are not academic. Mailing a "predicted high value" segment that is really a pile of stale, disengaged profiles is a fast way to damage sender reputation, which is the same mechanism behind the segment quietly burning your domain reputation.

Which catalogues can trust Predicted CLV, and which cannot?

The dividing line is repurchase behaviour, not revenue or list size. The more your customers buy on a regular interval, the more a frequency-based forecast has to work with.

Catalogue typeTypical repurchase behaviourPredicted CLV reliabilityBetter primary signal
Consumables (coffee, supplements, skincare)Regular reorders every few weeksHighPredicted CLV, replenishment interval
Apparel and accessoriesSeasonal, several times a yearModeratePredicted CLV blended with category and AOV
Considered durables (furniture, appliances, jewellery)Rare, often a single large orderLowHistoric CLV, first-order value, product tier
New or seasonal brand under 180 daysToo little history to modelNot availableHistoric CLV and engagement, until data matures

A brand can also sit in two rows at once. A furniture retailer with a candles-and-throws accessory line has a high-repeat tail hidden inside a low-repeat core, and a single account-wide Predicted CLV averages the two into a number that describes neither. Segment the catalogue first, then decide where the prediction earns trust.

What should you use instead when the model does not fit?

Fall back to signals you can defend, and reserve Predicted CLV for the parts of the catalogue that repurchase. Historic CLV is the honest anchor, because it is recorded fact rather than forecast, and for considered catalogues it usually predicts the future better than the model does. First-order value and the product tier purchased carry more information than frequency when frequency is close to one.

Klaviyo also supports a custom CLV definition, where you set an average order value, an average purchase frequency and an expected customer lifespan by hand. For a low-repeat catalogue those hand-set inputs, grounded in your own finance numbers, beat a frequency model fighting the wrong distribution. Where you do keep Predicted CLV, treat it as one input to a tier rather than the whole rule: pair it with an engagement condition and a recency guard so a mispriced forecast cannot, on its own, promote or suppress a customer.

What should you watch before segmenting or suppressing on Predicted CLV?

Watch three things, because each turns a modelling quirk into a revenue mistake.

First, the weekly retrain means the same profile can cross a threshold you built a segment on without any behaviour change, so a "Predicted CLV over X" tier will churn its own membership week to week. Set the boundary well clear of the crowd, and prefer bands over hard cutoffs. Second, decimals are not precision. Klaviyo reminds operators that predictions work best averaged over many customers and are not expected to be exact for any single individual, so a predicted 1.43 orders is a distribution, not a promise, and a per-person automation that reads it literally will misfire on the tails.

Third, never let a low prediction quietly gate a customer out of your best campaigns. On a considered catalogue the model systematically undervalues the year-old buyer who is, in fact, your ideal repeat prospect. If Predicted CLV feeds a suppression, audit what it removes before you trust it, and confirm you are not training the model on the very customers it keeps getting wrong.

Sources

M
Miloš Mitrović
Revenue Operations & AI Automation

Have a question or a project?

Whether it is about this post or a system you want built, I'm happy to talk.

Get in touch

404

Post not found. It may have been moved or the link is incorrect.

← Back to the blog
Ask AI About Me
Clicking an assistant copies the prompt and opens it: ready to run in ChatGPT, Perplexity, and Grok; in Claude, Gemini, or Copilot press Ctrl+V (Cmd+V on Mac) to paste. Use Copy prompt for any other AI. The assistant reads my site, so it needs web access.
Summarize with AI
ChatGPT, Perplexity, and Grok open with the prompt ready to run. Claude, Gemini, and Copilot open a chat with the prompt copied; press Ctrl+V (Cmd+V on Mac) to paste. The full text is included, so it works even without web access.