AI systems for Shopify brands doing $3M–$30M

You bill on a schedule. So does the leak.

Recharge, Skio, Smartrr, Stay AI, Loop — whichever platform charges your customers, it is still running most of the defaults it installed itself with. Every failure those defaults cause repeats on exactly the same cycle your revenue does. That is the bad news and the whole opportunity: the same decline codes, the same cancel reasons and the same reorder windows arrive every single month, which is the one condition under which it is worth building a system to read them.

  • 9%

    of recurring revenue lost to failed payments

  • 20–40%

    of subscription churn is involuntary — cards, not decisions

  • 25%

    of lapsed subscriptions are attributed to payment failure

Sources: Baremetrics; Paddle / ProfitWell; Stripe.

  • Recharge
  • Skio
  • Smartrr
  • Stay AI
  • Loop
  • Klaviyo
  • Shopify Plus

We know your numbers

Not yours specifically — these are category benchmarks with named sources, and the point of putting them here is that you can check your own against them before you believe anybody. Where we have not measured something, the row is an em dash.

Subscription churn, retention and payment failure — category benchmarks
MetricBenchmarkSource
Recurring revenue lost to failed payments~9%Baremetrics
Involuntary share of total churn20–40%Paddle / ProfitWell
Lapsed subscriptions attributed to payment failure25%Stripe
Monthly churn, B2C consumer goods6.5%Recurly
Monthly churn, DTC panel7.1%Recharge
— of which voluntary4.1%Recharge
— of which involuntary3.0%Recharge
12-month retention, monthly billing~28%Recurly / Recharge panels
12-month retention, annual billing~62%Recurly / Recharge panels
12-month retention, median~42%Recurly / Recharge panels
12-month retention, top quartile65%+Recurly / Recharge panels
Your own involuntary sharemetric to confirm

Payment figures from Baremetrics, Paddle / ProfitWell and Stripe. Monthly churn from the Recurly consumer goods churn benchmark and the Recharge DTC panel. The 12-month retention rows come from those same two panels, and we name both rather than pin a figure to one panel we cannot attribute it to precisely. None of this is a forecast for your account — replacing the last row with your own number is the whole point of the audit.

One point of monthly churn is worth compounding out before you decide what to fix first. The churn calculator does that arithmetic, and the benchmark pages carry the rest of the data with its sources.

Four things that break on a schedule

None of these are unusual and none of them are anybody’s fault. They are what a subscription programme looks like two years after it was set up by whoever was available that week. Each one is also a decision being made the same way every month by something that cannot see the data underneath it.

The retry ladder is a default nobody has touched

Every subscription platform ships with a retry schedule, and almost every brand is still running the one the app installed itself with. It is flat: the same number of attempts, on the same days, for the card that hit its limit on the 1st, the card that expired in March, and the card the bank flagged over an address mismatch.

In a one-off business a bad default costs you once. In a recurring one it bills again in thirty days, fails again for the same reason, and posts the same loss to the same line of the P&L — twelve times a year, compounding against a base that is also shrinking.

The information needed to do better is already in the payload. Every failed charge arrives carrying a decline code that says which of those three cards it is, and almost nobody reads it. A ladder that routes on the code rather than on the calendar behaves differently in each case: a soft decline retried on the schedule that code deserves, an expired card skipping the retries entirely and going straight to a one-tap update page, a hard decline handed to a person on the first failure rather than the fourth. Same platform, same subscribers, different system.

Cancel is the only door in the room

Three unopened bottles on a shelf is not a product failure. It is a cadence failure being expressed as a cancellation, because cancel was the only self-serve option on the page. The customer wanted the next one in March. They were offered leaving.

Households change quantity and cadence constantly — someone moves in, someone stops taking it, the last order is still half full. Pause, skip, resize, swap the SKU and change the interval all exist in your platform. Whether they are reachable in two taps, and whether the cancel flow offers them before it offers the exit, is a build decision nobody made.

A cancel flow worth the name reads the reason before it offers anything, then picks the door from it: “too much product” gets a cadence change, “too expensive” gets a smaller size, “I am travelling” gets a pause with a date on it. Discounting everybody is not a save strategy, it is a way of teaching your best subscribers to threaten to leave — and it is exactly what a flow with no logic in it does by default.

Reorder timing runs off the billing date, not the consumption curve

A 30-day supply at two capsules a day is a predictable reorder window. A 60-serving tub at one scoop a day is a two-month one. Both get billed on the same 30-day cycle because that is what the app defaulted to, and the second customer accumulates product until they pause — or until the shelf makes the decision for them.

The reorder prompt landing three days early beats it landing two weeks late, and both beat a prompt timed against a calendar that has nothing to do with how fast the product is actually used.

This is the one on the page that is genuinely a prediction problem rather than a configuration one. The fix is a model of consumption per SKU and per customer segment, built from your own order and reorder history, with the prompt, the education sequence and the billing cadence all timed against that curve instead of against the charge date. The platform cannot infer it, because the platform has never been told how much is in the tub.

The second order decides everything, and month three ends it

The first delivery is a trial. The second is the decision, and it is made in the gap between them — the weeks where nobody sends anything because the flows stop at the order confirmation.

Where the result is slow to show — a supplement, a skincare regimen, anything with a 60- to 90-day effect window — subscribers quit before the product has had time to work. That is a content and timing problem sitting inside a billing system, and it is the single most fixable cliff in a subscription business.

It is also the same consumption model as the row above, used for a different message: the education has to land while somebody still has product in the cupboard and no result yet, which is a date you can calculate rather than guess.

Three of the four are systems problems with a marketing symptom. They get solved in the retry ladder, the cancel-flow branch and the consumption model — not in the subject line.

What we build for subscription brands

Built in your accounts, documented as we go, and yours if we ever part ways. Most brands start with one of the first two and add from there.

  • Payment recovery The retry ladder rebuilt around what the decline code actually says rather than a flat schedule — which codes get retried, how many times, on which days, and which ones skip the retries and go straight to a card-update page that works in one tap on a phone. Pre-dunning before the charge that was always going to fail, and the account updater reconciled against the token Recharge, Skio, Smartrr or Stay AI is really billing, not just switched on at the processor.
  • Subscription retention A cancel flow that reads the reason before it offers anything, with more than one door behind it — pause, skip, change interval, change size, swap the SKU — plus the reason picker that tells you which of those doors people actually take, so the next round of save offers is built on your data rather than on a category assumption.
  • Lifecycle flows Onboarding through the second order, timed against a consumption curve modelled per SKU rather than against the billing date, and written for the customer who has not yet seen the result. Flows carry 41% of email revenue off about 5% of sends (Klaviyo benchmark data, 183,000+ brands) — this is the half of that number most subscription brands never build.
  • Fraud & chargebacks The other half of the money that leaves at the payment layer: screening rules tuned so they stop declining good subscribers, chargeback representment, and card-update coverage checked end to end rather than switched on and assumed.
  • Reporting & analytics Voluntary and involuntary churn split apart, cohort retention by billing cadence, and post-COGS profit per subscriber — sitting on a nightly job that reconciles Shopify against the subscription app against the 3PL and reports the rows that disagree, so the number arrives without anybody exporting a CSV.
  • Ops automation The manual work that grows with the subscriber count — address changes, skip requests, 3PL holds, failed-charge lists moved by hand between four tabs — rebuilt as workflows on n8n running on a VPS in your own hosting account, with your credentials and your backups.
  • Customer service AI The same six questions — where is it, skip the next one, change my address, pause me — answered from the live order and the live subscription record, and written back to the subscription so a skip actually skips the box instead of promising to. Refunds, discounts and any conversation where somebody has said “cancel” twice go to a person, by configuration.

None of these decide anything about money on their own. Refunds, credits, discounts, payment-method changes and cancellations that trigger a refund are assembled by a system and pressed by a person — in a business that bills every thirty days, an automated mistake does not end, it recurs.

What this looks like from the outside

We subscribe to real brands like a customer would, log every message with the day it landed, walk the cancel flow to the exit and check the reorder prompt against the date the product actually runs out. Then we publish it.

Client recovery figures

metric to confirm

We have no client results to publish yet, so there are none on this page. When there are, they will arrive with their baselines attached. The teardowns are public and unaffiliated in the meantime — the same analysis, run from the outside on brands who did not ask for it.

Read the teardowns →

Questions from subscription brands

We’re on Recharge. Do you work inside it, or do you want to move us?

Inside it. Recharge, Skio, Smartrr, Stay AI and Loop are all platforms we build in, and we would rather fix the retry ladder and the cancel flow you have than spend eight weeks migrating them somewhere else. A migration is occasionally the right answer, but it is a decision that should come out of an audit — not out of a preference we brought with us.

Where does the AI actually sit in this? Which part is a model?

Less of it than the category implies, and we would rather say so. The retry ladder, the dunning sequence, the pause and skip logic and the reconciliation job are deterministic workflows — rules are cheaper per run, testable, and incapable of inventing an answer, so anything a rule can do, a rule does. A model earns its place in three places: reading a free-text cancel reason or an inbound ticket, where the input is a customer’s actual sentence; modelling consumption per SKU, where the answer is a prediction rather than a lookup; and drafting language a person then approves. The agents on top — order status, skip and address changes — are models with tools, and each tool is one we wrote, scoped and logged.

Will an agent be talking to our subscribers?

Only inside a boundary we write down together before anything launches, and only after it has run in draft mode long enough for the logs to justify it: the agent drafts, your team sends, and autonomy is granted one intent at a time. Anything touching a refund, a discount, a payment method or a cancellation is assembled by the agent and pressed by a person. That is a deliberate boundary rather than a technical limit — billing is the one place where an automated mistake compounds every month instead of ending.

How much of our churn is actually involuntary?

Usually more than people expect. Across subscription businesses, involuntary churn accounts for somewhere between 20% and 40% of total churn (Paddle / ProfitWell), and Stripe attributes about 25% of lapsed subscriptions to payment failure. The audit splits your own cancellations so you are working from your number rather than the category’s.

We’re mid-migration between subscription platforms. Is this a bad time?

It is usually the best time. Retry logic, dunning sequences, cancel-flow branches and reorder timing all have to be rebuilt on the far side of a migration anyway, and most migrations carry the old defaults across untouched because the deadline was the storefront, not the billing logic. Building it once, deliberately, during the move costs less than rebuilding it twice.

We bill annually as well as monthly. Does the same work apply?

Yes, with different emphasis. Annual billing retains far better — roughly 62% at twelve months versus about 28% for monthly (Recurly / Recharge panels) — but it concentrates every failure into one charge a year, so a single declined renewal is a whole customer rather than a month. Annual programmes need pre-dunning and card-expiry warning more than monthly ones do, not less.

Our email agency owns Klaviyo. Does this conflict with them?

No. Most of this is billing infrastructure rather than campaign marketing — retry timing, decline-code routing, cancel-flow logic, the card-update path — and it is usually the piece nobody at a retention agency owns. Where the work does touch Klaviyo, we build in your account, document what we changed, and leave the campaign calendar to them.

How fast does any of this show up in the numbers?

Recovery starts on the first billing cycle after launch, because the charges are already failing on a schedule. Reading the result properly takes two full cycles — one to recover, and one to confirm the recovered subscribers stayed. Cancel-flow and reorder-timing changes take a full purchase cycle before the numbers mean anything, and a consumption model needs enough reorder history behind it to be worth trusting — if you do not have that yet, we will say so rather than model it anyway.

Do we need to change payment processors?

Almost never. Nearly everything we change sits above the processor — retry timing, decline-code routing, the dunning sequence and the card-update path. If something at the processor level is genuinely costing you money, we say so in the audit and show the working.

We do $4M with one ops person. Are we too small for this?

No — that is the middle of who we build for. The published floor is $3M+ annual revenue, on Shopify Plus or running a paid subscription platform. Below that, a five-figure project is a large share of a year’s profit and the honest advice is to spend it elsewhere, which is what we will tell you on the first call.

Find out what you’re losing.

Before you commit to anything, we tell you exactly what you’re losing and what it costs to stop it. Two weeks. Fixed fee. Credited in full against any build you go ahead with.

Fee
$1,500–$3,000, fixed
Duration
Two weeks
Credited
In full, against any build
You supply
Read access + one 45-minute call