What do good Klaviyo benchmarks look like at $3M-$30M revenue?
Good Klaviyo benchmarks are a range built from your own list size, flow count and send history, not a single number borrowed from a vendor report. Klaviyo states that automated flows generate 41% of total email revenue across its aggregated base of 183,000+ brands (vendor-reported), and that figure gets quoted as a target by almost every agency deck in the category. It isn’t wrong. It’s just not segmented by the two variables that actually move it: how many flows a brand runs, and how large its list is.
A brand at $3M with three flows and a brand at $30M with fourteen flows both sit inside that 183,000-brand sample. Averaging across them produces a number neither brand should treat as their own target. The useful version of a benchmark answers a narrower question: given your flow count and list size, what should your flow revenue share look like in six months if the programme is healthy?
What does Klaviyo’s own benchmark data say?
Klaviyo’s public benchmark reporting gives one cleared figure worth quoting directly, and it stops short of the granularity most operators actually need.
| Metric | Figure | Source status |
|---|---|---|
| Share of email revenue from automated flows | 41% | Vendor-reported (Klaviyo, 183,000+ brands) |
| Open rate by list size | Not published at this granularity | Metric to confirm in your own account |
| Click rate by flow type | Not published at this granularity | Metric to confirm in your own account |
| Revenue per recipient by flow type | Not published at this granularity | Metric to confirm in your own account |
What to take from this table: the only figure Klaviyo publishes with enough weight to quote is the aggregate flow-revenue share, and it comes with no breakdown by revenue tier, vertical or flow count, which means every other row has to be built from your own account rather than pulled from a report.
Why do vendor averages mislead brands above $3M?
An average built from 183,000 brands pulls in stores selling $50,000 a year alongside stores selling $50M, and mixes brands running one welcome flow with brands running a full lifecycle programme. A $3M-$30M brand sits in the part of that distribution where the gap between “does the basics” and “runs the full set” is largest, so the average lands nowhere close to either group.
There’s a second distortion specific to this revenue band. Brands at this size often have list sizes large enough to carry meaningful cold and lapsed segments, which drags down campaign engagement rates and inflates the flow revenue share for reasons that have nothing to do with flow quality: triggered sends to recently engaged profiles will always beat blast sends to a five-year-old list, regardless of how good either programme is. Reading a rising flow-revenue-share number as “flows got better” without checking whether campaign volume or list hygiene changed first is the most common misread of this metric.
How do open, click and revenue rates change by list size?
No cleared industry figure exists at the resolution of “open rate for a list this size,” and any number presented as one should be treated as invented. The honest answer is structural: as a list grows, it accumulates more cold and unengaged profiles unless suppression is actively managed, which mechanically lowers blended open and click rates even when the engaged segment’s behaviour hasn’t changed.
The method that works without a borrowed number: segment your list into engaged (opened or clicked in the last 90 days), lapsed (no engagement in 90-180 days) and cold (no engagement beyond 180 days), then pull open and click rates for each segment separately. A healthy engaged-segment open rate that’s flat quarter over quarter, alongside a growing cold segment, tells you the blended number is dropping for structural reasons, not because your flows got worse. That distinction changes what you fix: suppression and list hygiene, not creative or subject lines.
Which benchmark actually predicts flow revenue?
Flow count and flow coverage predict flow revenue far more reliably than any engagement-rate benchmark, because flow revenue is a function of how many buying moments are automated, not how good any single flow’s copy is.
| Flow present | Typical role | Revenue mechanism |
|---|---|---|
| Welcome series | First-purchase conversion | Captures intent at the point of highest interest |
| Browse abandonment | Recovers product-page exits | Reaches shoppers before cart, not after |
| Cart/checkout abandonment | Recovers near-purchase drop-off | Highest intent, shortest window to act |
| Post-purchase | Sets expectations, opens cross-sell | Reaches customers at their most engaged |
| Replenishment | Time-based repurchase prompt | Only applies to consumable or wear-out products |
| Winback | Reactivates lapsed customers | Recovers revenue that would otherwise churn silently |
What to take from this table: a brand missing browse abandonment and winback is leaving two structurally different revenue mechanisms on the table (one that catches interest before commitment, one that recovers customers after they’ve gone quiet), and no amount of subject-line testing on the flows you already have will replace either. If you want to see what a missing flow is worth in your own numbers before building it, run your own figures through the flow revenue calculator at /tools/flow-revenue-calculator rather than estimating.
Where do most $3M-$30M teams sit against these numbers?
Teams in this band commonly run a handful of flows, usually welcome, cart abandonment and post-purchase, with browse abandonment, winback and replenishment either missing or built once and never revisited. That’s not a criticism of the team; it’s usually a sequencing decision made when the list was smaller and the missing flows weren’t worth the build time yet. The problem is that the same three-flow set often stays in place well past the point where the list and order volume justify more.
The tell is a flow-revenue share that plateaus rather than grows as the list scales. As an illustrative case: a list growing at a steady rate with a static flow set will show flow revenue growing roughly in line with the list, while the same list growth rate paired with an expanding flow set (segmented further, new triggers added) should show flow revenue growing faster than the list, because each new flow captures revenue the old set was structurally unable to reach.
What breaks when a team chases the wrong benchmark?
The most common failure is optimising the number instead of the mechanism. A team that sees flow revenue share below 41% and responds by cutting campaign sends to shrink the denominator will hit the target number without adding a single dollar of flow revenue: the share moves, the business doesn’t. That’s a vanity fix, and it usually costs real campaign revenue to buy a benchmark that looked better on a slide.
Treating open rate as the metric to chase inside flows specifically is a second failure mode. Open rate is a proxy for deliverability and subject-line relevance, not for revenue. A flow can carry a strong open rate and a weak click-to-purchase path, or a modest open rate against a highly engaged segment that converts at volume. Chasing open rate in isolation leads teams to rewrite subject lines quarter after quarter while the actual leak (a weak offer, a broken link, a missing segment split) goes untouched.
A third failure mode is comparing this quarter’s numbers to a competitor’s public case study rather than to the brand’s own trailing average. Case studies report the best quarter, not the typical one, and almost never disclose list size, flow count or suppression practice, which makes the comparison meaningless even when the headline number looks close.
Who this data does not apply to
This benchmark framing assumes a brand already running Klaviyo with a real transactional history — enough order volume to make flow segmentation meaningful, and enough list size for the engaged/lapsed/cold split to carry statistical weight. A brand under roughly $3M in revenue, or one still on a starter ESP without flow branching, won’t have the order or list volume to make most of these comparisons stable quarter to quarter; small-sample swings will look like trends that aren’t real. That’s not a reason to ignore benchmarks, just a reason to hold them more loosely until the account has more history behind it. Brands scaling past that point are covered in more detail at /for/scaling-brands.
A benchmark, vendor-reported or otherwise, is a starting point for a conversation with your own trailing data, not a number to hit and stop; it doesn’t replace a working knowledge of your own account’s history.
Most teams that miss on these benchmarks aren’t missing on execution inside individual flows: they’re missing flows entirely, or running the right flows against the wrong segments. That’s a lifecycle flows problem before it’s a copywriting or deliverability one, and it’s the reason a benchmark review should start with flow coverage and segmentation, which is what our lifecycle flows work is built around.
Sources
- Klaviyo, 183,000+ brands: 41% of email revenue from automated flows (vendor-reported), quoted as Klaviyo’s own aggregated benchmark figure and not independently verified in this article.