Fraud screening, chargebacks & card updates

Keep the revenue your payment layer loses.

Every order your filter blocks is a classification, and yours has only ever been tuned in one direction. We rescore it against orders you actually shipped, type and assemble every dispute the moment it arrives, and refresh card credentials before the charge fails.

  • 25%

    of lapsed subscriptions are attributed to payment failure

  • 20–40%

    of all subscription churn is involuntary — cards, not customers

Sources: Stripe; Paddle / ProfitWell. We have no independent figure for false-decline, chargeback or account-updater rates in this category, and we do not print one — those come out of your own processor during the audit.

The problem

Three mechanisms lose money at the payment layer, and most brands treat them as unrelated problems with different owners. A false decline is a good customer refused at the door. A chargeback is revenue you already recognised, taken back weeks later, with a fee and a deadline you probably missed. An expired card is a subscriber who never chose to leave. Same layer, same month, three separate exits.

This page is one half of a pair. Payment recovery covers the charge that failed — retry timing, decline-code routing and the dunning sequence after it. This one covers what that page does not: the screening decision before the charge, the dispute that arrives after it, and the card credentials that go stale in between. Halves of one problem, usually built together — but different work, so different pages.

For a subscription brand each of these compounds, because the charge repeats. A one-off store loses one order to a false decline; you lose that order and every one after it, the whole remaining lifetime of a customer whose card was fine. A chargeback is not a one-box problem either: it is the box, the fee, the ratio that fee counts against, and the subscription that ends alongside it.

The fraud filter is refusing customers you want

A screening decision is a classification with two ways to be wrong, and only one of them is ever counted. A fraudulent order that gets through lands in a report with a fee, a reason code and somebody’s name attached. A good customer refused at the door leaves no record at all — no ticket, no dispute, no line item, usually not even a complaint, because they assume the store is broken and buy elsewhere. One error is measured and the other is invisible, so the threshold only ever moves one way. Worse, the two errors are not even priced alike here: a fraudulent order costs you one box and a fee, while a refused subscriber costs you every renewal they would have paid for. The expensive mistake is the one nothing counts. That asymmetry, not the tool, is why the filter ends up refusing people you want.

So rules accumulate — block this country, block a mismatched AVS, block the third order from one address in an hour — each defensible when written, none ever removed. Meanwhile a renewal looks odd in ways a first order does not: same card, same amount, same address, month after month, and then a gift bag shipped to a different state in December. The customer never sees a fraud message. They see a decline, and a decline reads as a broken store.

The information that would settle it is already on the account and no rule is reading it: eleven delivered boxes, two years of tenure, the same card, a support history with no disputes in it. A first-time stranger paying with a new card has none of that. Those are two different populations and they need two different decisions — which is a modelling problem, and it is the one thing a rule list written a country at a time cannot express.

The chargeback nobody contested

Disputes arrive with a clock running and land in an inbox with no name on it. The evidence that wins one exists — the subscription agreement they accepted, the delivery confirmation, the login and IP history, the eleven boxes they took without complaint — but it sits in five systems and takes an hour to assemble, so the case is let go by default. That is an administrative loss, not a commercial one. And plenty of these are not fraud: “I forgot I was subscribed” arrives as a dispute rather than a cancellation, because the dispute form was easier to find than your cancel flow.

A dispute is a classification with a deadline: what kind of case is this, what evidence would answer it, does that evidence exist for this order. All three questions have answers a system can produce in seconds, out of the subscription record, the carrier scan and the message log. That is the part worth automating — the gathering, the typing and the clock — because the hour it takes by hand is the entire reason winnable cases go unanswered. What the issuer reads is still written and sent by a person.

The card that was always going to expire

Every stored card has an expiry date, and on a repeating charge that date is a scheduled failure with a known cause. Cards also get reissued after a breach or renumbered when a bank switches processor. Account and card updater services exist at Stripe and the major processors precisely for this, and are usually either switched off or switched on and never reconciled — so the refreshed credential sits at the processor while Recharge, Skio, Smartrr or Stay AI keeps billing the token it already had.

Where the decision stops being automatic

The system scores an order, types a dispute, assembles its evidence, watches the deadline and reconciles a refreshed card. It does not refund, credit, cancel or re-price anything, and it does not file a case on its own. Two calls in particular stay with a person: blocking a returning subscriber whose history says they are real, and answering a dispute where the honest read is that the customer has a point. Both are cheap for a human to make and expensive for a system to get wrong — a false decline is a customer lost silently, and an aggressive representment against a genuine complaint buys a fee and loses the relationship. Where a wrong answer costs more than a human minute, a human spends the minute.

What we build

Eight pieces across the three mechanisms, built in your accounts and documented as they go. Most brands need all three mechanisms addressed; few need all eight pieces on day one.

  • Screening, measured like a classifier Every active rule in Stripe Radar, Signifyd or Riskified listed, dated and replayed against twelve months of orders you actually shipped — scored the way any classifier is scored, with the false positives counted rather than assumed. Rules that have never stopped a fraudulent order and have only ever stopped customers come off.
  • A different score for renewals A subscription renewal is not a first order and should not be classified like one. A returning subscriber carries evidence a stranger with a new card cannot fake — eleven delivered boxes, the same amount on the same card each month, an address that has not moved — and gets a path through screening that reads it.
  • Decline reporting, split by cause Issuer declines separated from screening blocks, because they are two different problems with two different owners. Without that split, the retry decision on the other half of this pair is being made blind — it cannot learn from a failure your own filter caused.
  • Chargeback evidence packs One assembled template per reason code — product not received, unauthorised, subscription cancelled — pulling the subscription agreement the customer accepted, delivery confirmation, IP and login history, message engagement and every prior box they took without complaint into a single submission. The classification and the assembly are the automated half: the case is typed by reason code and the evidence pulled out of five systems in seconds. What goes to the issuer is still read by a person.
  • Dispute routing and deadlines Every new dispute is typed by reason code, matched against the evidence that actually exists for that order, and lands somewhere with a name on it and a countdown attached rather than in a shared inbox. A missed deadline is the most common reason a winnable case is lost, and it is an administrative failure rather than a commercial one.
  • Card and account updater, wired end to end The updater service at Stripe or your processor switched on, then reconciled against the subscription platform — so the refreshed credential actually reaches the token Recharge, Skio, Smartrr or Stay AI bills against, instead of sitting one system away from it.
  • Expiry inventory Every stored card with an expiry date inside the next quarter, listed and watched, so a failure that was always going to happen is visible before it happens. What gets said to those subscribers is dunning, and that lives on the other half of this pair.
  • A measurement baseline Your false-decline, chargeback, representment and updater numbers, established before we change anything — because for these four there is no category benchmark we would be willing to hand you as though it were yours, and because a classifier nobody has measured is a classifier nobody can tune.

How the build runs

Four weeks from access to launch, on a fixed scope and a fixed price. The first week is measurement, because nothing here should be changed against a benchmark borrowed from somebody else’s account.

  1. 01

    Baseline & instrumentation

    Twelve months of declines pulled and split into issuer declines and screening blocks. Every dispute with its reason code, evidence and outcome. Every stored card with an expiry inside the next quarter. You get your own four numbers here, because the published ones for this half are not worth quoting — and because a screening decision has nothing to learn from until its past calls have outcomes attached to them.

    Week 1

  2. 02

    Screening rules, rebuilt

    Each rule in Stripe Radar, Signifyd or Riskified replayed against orders you already shipped, so you can see what it caught and what it refused — a confusion matrix rather than an opinion. Then a renewal path that scores a returning subscriber on tenure and delivery history rather than on the signals a first-time stranger is judged by, and thresholds set against the revenue on each side of the line instead of against fraud losses alone.

    Week 1–2

  3. 03

    Representment system

    Evidence packs built per reason code, wired to the systems the evidence lives in, so a case is typed and assembled in seconds and takes minutes to file instead of an hour. Routing, ownership and deadline alerts on top, because the cases lost by default outnumber the cases lost on the merits. A person still reads the pack before it is filed — the automation is the gathering, not the argument.

    Week 2–3

  4. 04

    Card & account updater

    Updater enrolment at Stripe or your processor, then the part almost everyone skips: reconciling the refreshed credential back into Recharge, Skio, Smartrr or Stay AI, and proving with a test renewal that the new number is the one being charged.

    Week 3

  5. 05

    QA & launch

    Screening changes run in shadow first — scoring every order and deciding none of them — so what they would have blocked can be read against what actually shipped and got paid for. Evidence packs filed on real closed cases to check the fields resolve, and updater reconciliation verified on a test subscription. Nothing goes live untested — a screening rule that refuses good orders is more expensive than no rule.

    Week 4

  6. 06

    Measure & iterate

    The four baseline numbers re-read every month against week one. Screening drifts as your traffic changes — a new market, a bigger gifting season, a different acquisition channel — so thresholds are re-read rather than left where they were set. Rules that stop earning their place come off, evidence packs get rewritten against the reason codes you actually receive, and the updater reconciliation is checked rather than assumed.

    Ongoing

What the published numbers are worth

There is one figure at the payment layer that several parties have published, and they do not agree with each other. It belongs to the other half of this pair, but it is worth looking at here, because how it was produced tells you what to do with every percentage you get quoted about your own payment layer — including ours.

Recovery rate on failing invoices — four published figures
Reported byRecovery rateAttribution
RecurlyUpwards of 80%Vendor-reported
Chargebee50–70%, with automationVendor-reported
Stripe smart retries~38%, with no emailsVendor-reported
Slicker / Baremetrics~47.6% medianIndependent

Three of these four rows are a vendor describing its own product, on its own definition of a recovered invoice, across its own customers. The fourth is the only independent median we could find. Read together, the vendor figures disagree with each other by more than a factor of two — about 38% at one end and upwards of 80% at the other — and the highest of them sits roughly 1.7× above the independent median. That multiple is our own arithmetic on the four published figures rather than a finding. The lesson is not that anybody is lying; it is that a number produced by whoever benefits from it is not a forecast for your account.

Now notice what is missing from that table. There is no row for false declines, no row for chargeback rates, no row for representment win rates and no row for account-updater catch rates — the four figures this page would most like to quote you. We could not find an independent source for any of them that we would stand behind, and the vendor numbers that do circulate have the same problem the table above demonstrates, only without an independent median sitting next to them for scale.

So we leave them as em dashes and produce yours instead. That is not a gap in the page; it is the argument the page is making.

The four figures this page will not invent
MetricOur published figureWhere yours comes from
False-decline rateMetric to confirm — your screening tool’s block log, replayed against shipped orders
Chargeback rate and reason mixMetric to confirm — your processor’s dispute record, twelve months
Representment win rateMetric to confirm — measured on your own cases after launch
Account-updater catch rateMetric to confirm — your processor, reconciled to the subscription platform

Each of these stays an em dash until it has been measured on your account. Quoting a vendor’s figure as though it were yours is how a brand ends up budgeting against a percentage that was never about its business. The audit produces all four in about a week, and they become the baseline everything afterwards is read against.

What it costs

Published ranges, because hidden pricing costs more leads than it protects. Where you land inside a range depends on how many processors, screening tools and subscription platforms are in play, not on how much we think you can pay.

Revenue Recovery Audit — covers both halves of the payment layer

$1,500–$3,000

Fraud, chargeback and card-update build

$4,000–$10,000

Built together with payment recovery, as one payment-layer project

Scoped as one build

Ongoing operation — cases filed, rules tuned, updater reconciled

By arrangement

The audit is credited in full against any build you go ahead with, and it covers both halves of the payment layer — splitting the pages does not mean splitting the fee. For scale on the tooling side of this problem: recovery and dunning apps in this category are commonly published at $300–$2,000 a month, or 5–15% of the revenue they recover. That is a market range rather than ours, and it is worth knowing, because a percentage-of-recovery fee is a standing tax on money you would rather keep — the point of owning the system is that it stops being one.

Questions

Why is this a separate page from payment recovery?

Because it is different work in different systems, even though the money leaves from the same place. Payment recovery is about the charge that failed: when to retry it, how often, and what message follows — a timing decision made per charge. This page is about the classification made before the charge, the dispute that arrives weeks after it, and the card credentials that go stale in between. We usually build them together, and one audit covers both.

What is our false-decline rate?

We do not know, and neither does anyone quoting you an industry average for it. The published false-decline figures are almost all produced by vendors selling screening, on their own definitions, and we could not find an independent benchmark for subscription brands on Shopify that we would be willing to print. So we take yours instead: the screening tool’s block log replayed against twelve months of orders, matched against which of those refused customers came back and which never did. That is the first thing the audit does.

Do we have to replace Stripe Radar, Signifyd or Riskified?

Almost never. The tool is rarely the problem — the rule set is, because rule sets drift in one direction. Each rule is defensible on the day it is written, nobody is measured on the revenue it refuses, and so none of them ever come off. All three tools already score orders; what is missing is a measurement of how well that score does on your traffic, and a second decision for renewals that reads tenure and delivery history instead of the signals a first-time stranger is judged by. We work inside whatever you already run. If the tool genuinely cannot express that distinction, we say so and show the working.

Do you file the chargeback responses, or do we?

Either. The build hands your team the evidence packs, the routing and the deadline alerts, so filing a case takes minutes rather than an hour. If you would rather not own it at all, filing can sit inside an ongoing arrangement. What we will not do is quote you a win rate to sign against.

What win rate should we expect on representment?

We will not give you one, and it is worth being careful with anyone who does before looking at your account. Outcomes turn on the reason code, the issuer, the evidence you can actually produce and how much of your dispute volume is friendly fraud rather than real fraud. What we can describe is the mechanism: a case answered inside the deadline, with the subscription agreement, the delivery confirmation and the customer’s own order history attached, is a materially different case from one answered late or not at all. We measure yours after launch, and we publish nothing until there is something honest to publish.

What is an account updater, and do we already have one?

It is the service that tells you a stored card has been reissued with a new number or a new expiry date, so the next renewal charges the new credential instead of failing against the old one. Stripe offers it and so do the major processors. Most brands we look at are in one of three states: it is switched off, it is switched on and nobody has checked it, or it is switched on at the processor and never reconciled with the subscription platform — so the card updates in one system while Recharge, Skio, Smartrr or Stay AI keeps billing the token it already had. We establish which of the three you are in week one.

We only see a handful of chargebacks a month. Is this worth doing?

Possibly not on the chargeback half alone, and we will say so rather than sell you the build anyway. But these three mechanisms rarely arrive one at a time, and the one quietly costing a subscription brand the most is usually the least visible: cards going stale against a charge that repeats every month. If the numbers do not justify a build, the audit says so and points at the leak that does. If we cannot identify recoverable revenue worth more than the audit fee, we refund it.

Will loosening the fraud rules increase our fraud losses?

It can, and pretending otherwise would be dishonest. Moving a threshold trades one error for the other, which is exactly why the work starts with measurement rather than with loosening anything — every rule replayed against orders you already shipped, so you can see what each one caught and what each one refused, then run in shadow before it decides anything live. Some rules come off, some get tightened, and a few get replaced by a renewal path that did not exist before. The goal is not less screening. It is a screening decision priced against the revenue on both sides of the line rather than against fraud losses alone.

What does the system decide by itself?

It scores orders, types each dispute by reason code, assembles the evidence for it, tracks the deadline and reconciles a refreshed card back into the subscription platform. It does not refund, credit, cancel or re-price anything, and it does not file a case on its own. Two decisions in particular stay with a person: blocking a returning subscriber whose history says they are real, and answering a dispute where the customer plainly has a point. Both take a human a minute and cost real money when a system gets them wrong.

Find out what you’re losing.

Before you commit to anything, we tell you exactly what you’re losing and what it costs to stop it. Two weeks. Fixed fee. Credited in full against any build you go ahead with.

Fee
$1,500–$3,000, fixed
Duration
Two weeks
Credited
In full, against any build
You supply
Read access + one 45-minute call