AI systems for Shopify brands doing $3M–$30M
The stack that got you here stopped fitting.
Spreadsheets, a shared inbox and eleven tools nobody owns. That worked at $1M. Somewhere past $3M it became the reason the team is busy and the margin isn’t — and the default answer is another hire, when most of what that hire would do is read one screen and retype it into another.
-
8–12
marketing tools the typical $15M+ brand runs, rarely reconciled
-
4–8%
EBITDA at $5M–$10M in revenue — the margin every extra hire comes out of
Tool count: Pointerflow category review. Margin band: DTC operator benchmarks compiled in our ICP research. Both are category patterns — check them against your own P&L.
- Shopify Plus
- A subscription platform
- Klaviyo
- A paid helpdesk
- A 3PL
- Two attribution dashboards
- Spreadsheets
We know your numbers
The break is predictable, which is the only good news in it. These are the milestones it tends to arrive at and the margin it arrives inside — useful for sizing a decision, useless as a forecast. The last two rows are yours, and they stay em dashes until they have been measured.
| What we look at | Where the line sits | Basis |
|---|---|---|
| Crossing roughly $1M | Operational complexity passes what a founder and a small team can hold in their heads | Pointerflow ICP research — a pattern, not a measurement |
| First Head of Operations or Head of Retention posted | The inflection point: the work has outgrown the founder | Pointerflow ICP research — a pattern, not a measurement |
| EBITDA at $5M–$10M revenue | 4–8% | DTC operator benchmarks compiled in our ICP research |
| EBITDA at $10M–$30M revenue | 8–14% | DTC operator benchmarks compiled in our ICP research |
| Marketing spend, DTC median | 13.3% of revenue | DTC operator benchmarks compiled in our ICP research |
| Marketing tools running at $15M+ | 8–12, rarely reconciled | Pointerflow category review — a category estimate |
| Your own post-COGS profit by channel | — | metric to confirm |
| Hours a week your team spends reconciling | — | metric to confirm |
Every row here is a category pattern to argue with rather than a number to plan against. The margin bands and the median spend figure are DTC operator benchmarks we compiled while defining who this work is for; the tool count is our own category review. Replacing the last two rows with your own numbers is the whole point of the audit.
Six things that break somewhere past $3M
Not one dramatic failure. Six quiet ones that arrive together, none of which is anybody’s fault, and all of which get answered with headcount because headcount is the only lever anybody has been given.
The stack was assembled, not designed
Every tool in it solved one week’s problem. The subscription platform, the helpdesk, the reviews widget, the two attribution dashboards that disagree, the fourth thing somebody trialled and never switched off.
None of them was chosen against the others, so nothing reconciles — and the work of making them agree gets done by people, every week, indefinitely. That labour never appears as a line item anywhere, which is why it never gets challenged.
The flows were built once, by someone who has left
A welcome series written two years ago against a product line you have since changed. A winback firing on a rule nobody can explain. Three flows that are live and send to almost nobody.
None of it is broken in a way that raises an alarm, which is exactly why it has survived. Nobody owns them because ownership was never assigned to a role — only to whoever happened to set them up.
The shared inbox is the system of record
Where an order actually is. Whether a refund was agreed. Which subscriber was promised a free replacement, and which one was told the reorder would land before the current 30-day bottle ran out.
All of it is real operating information, none of it is in a system, and it works fine right up until the person who reads that inbox takes a holiday.
Nobody can answer the profit question
Ask for post-COGS profit by channel and you get a pause, then a spreadsheet with a caveat attached. A brand past $15M is typically running 8–12 marketing tools, each reporting a version of revenue it attributes to itself.
Not one of them carries the cost of goods, the real outbound shipping charge off the 3PL invoice, or the discount that was applied at checkout. So the number everyone plans against is the one nobody can defend.
The default answer is another hire
The fix becomes a coordinator, then a second one, then a manager to run the coordinators. Three people end up doing what one script would do at three in the morning — and the salary line recurs every year while the script is paid for once.
Some of that headcount is genuinely necessary: judgement, supplier relationships, the decisions that need a person. The copy-paste between two dashboards is not, and the two rarely get separated before the offer goes out.
And the margin it comes out of is thin
DTC EBITDA at $5M–$10M in revenue typically runs 4–8%. At 6% on $8M that is $480,000 of profit for the year — arithmetic on a benchmark band rather than your number — set against a median marketing spend of 13.3% of revenue.
In a shape like that, an operating cost you carry because one tool doesn’t talk to another is not a rounding error. It is a visible share of what the year was for.
The first four are systems problems wearing an operations symptom, and they get solved in the data model, the workflow and the runbook rather than in a job description. The last two are what happens when nobody names them as systems problems in time.
What we build for brands at this stage
Seven pieces, in roughly this order. Before any of them, every manual job gets named, timed and costed, then sorted into three piles: what a deterministic rule should run, what an agent should draft for a person to send, and what nobody should automate. The third pile is real, and we argue for it. Everything is built in your accounts, documented as we go, and yours if we ever part ways.
- Reporting layer A warehouse in your own cloud account carrying order-level margin: cost per variant, inbound freight, packaging, the real outbound shipping charge off the 3PL invoice, payment fees and the discount applied. Post-COGS profit by channel becomes a column rather than a quarterly exercise. This goes first, because everything after it gets argued from numbers instead of opinions.
- Definitions doc What counts as a new customer, which discount codes are cost of goods, when a refund lands, how a subscription rebill is classified. Boring, and the reason two people stop pulling different versions of the same number.
- Flows, rebuilt Every live flow mapped, the dead ones switched off with the reason written down, and the ones actually carrying revenue rebuilt against the product line you sell now. Ownership assigned to a role rather than to whoever set them up.
- Payment layer Retry ladder, pre-dunning, dunning sequence and a card-update page that works in one tap on a phone. If you bill on a schedule, this is the failure that repeats every month until somebody owns it.
- Inbox jobs Refund checks, order lookups, 3PL exception handling and the recurring spreadsheet exports, moved onto self-hosted n8n in your own infrastructure. Each one written down as a workflow before it is built, so you can see exactly what is being removed and what is left for a person.
- Support triage The same handful of questions resolved without a human, and everything else routed to one with the order, the subscription state and the shipment already attached to the ticket.
- Agents & runbooks The jobs that need judgement on an unstructured input — a ticket, a supplier’s PDF, a review, a description that has to be written rather than looked up — built as agents on n8n in your own infrastructure, each running in draft mode until its own logs earn it the next permission. Then every workflow documented, in your accounts, with a written procedure for changing it. Your first ops hire inherits a system instead of a folder of tribal knowledge and three people to ask.
Not all seven are worth doing at your size, and the audit says plainly which ones aren’t. What gets fixed before the ops hire and what gets left for them is a sequencing decision, and it is part of what you are buying.
What this looks like from the outside
More of this is visible from outside than you would expect. We buy the product like a customer, log every message with the day it landed, walk the cancel flow to the exit and check the reorder prompt against the date the product actually runs out. Where a system should be and a person is instead, it shows.
Client results at this stage
—
metric to confirm
We have nothing of our own to publish yet, so there is no number here. When there is, it will arrive with the baseline it was measured against. Until then the teardowns are unaffiliated: the same analysis, run on brands who did not ask for it.
Read the teardowns →Where to go next
Questions from brands at this stage
We just posted a Head of Operations role. Should we wait until they start?
No — do it so the role is worth taking. A first ops hire who arrives to an undocumented stack spends two quarters reverse-engineering it before they can improve anything. One who arrives to a reporting layer, a set of runbooks and the manual jobs already automated spends that time running the business instead. We would rather build the thing they inherit than compete with them for it.
Is this instead of hiring?
Sometimes, and we will tell you when it isn’t. Judgement work, supplier relationships and anything that needs a person deciding are not automation problems. The copy-paste between two dashboards is. The audit separates the two and puts a number on the second, so the hiring decision gets made against evidence rather than against how busy everybody feels.
Is any of this actually AI, or is it just automation?
Both, and the split is the interesting part. Most of what we build at this stage is deterministic: a rule that reconciles two exports, routes an exception, retries a declined card on a schedule the decline code chose, or posts the same number to Slack every morning. A rule is cheaper per run, testable, and incapable of inventing an answer. A model earns its place where the input is genuinely unstructured — a customer’s actual sentence, a supplier’s PDF, a product description that has to be written rather than looked up. Anything that decides money is drafted by the system and pressed by a person. If we cannot tell you which of those three a given job is, we have not finished the triage.
Should we wait for the models to get better before building any of this?
No, because the models are not the bottleneck at your size. What kills an AI project here is the layer underneath it: eight to twelve tools that do not agree on what a customer is, a shared inbox holding operating facts no system can read, and no reconciled profit number to judge any of it against. That is exactly why the reporting layer goes first on this page. An agent pointed at data nobody trusts produces confident answers nobody should act on — and it produces them faster than a person could check them.
We’re at $2.5M. Are we too early?
Probably, and we would rather say so than take the money. The published floor is $3M+ annual revenue, on Shopify Plus or running a paid subscription platform. Below it, a five-figure build is a large share of a year’s profit at the margins this category runs on, and the honest advice is to spend it on the single leak costing you most rather than on a system. Ask us which one on a call and we will tell you without a proposal attached.
Do we have to rip out our tools?
Rarely. Most of what breaks at this size is the absence of a layer underneath the tools rather than the tools themselves. We build the warehouse and the automation in your own accounts and leave the front ends alone. Where something genuinely has to go — usually the second attribution dashboard nobody trusts — we say so and show the arithmetic.
How long does it take?
The audit is two weeks at $1,500–$3,000, fixed, and credited in full against any build. Builds run in four-week blocks at a fixed price per block. The reporting layer normally goes first, because every decision after it is easier to argue when the numbers agree.
Who maintains it once you’re gone?
Your team, which is the point. Everything is built in your accounts, on tooling you pay the vendor for directly, with a runbook per workflow. Self-hosted n8n is the automation layer precisely because it does not hold the business hostage to a per-task price, or to us.
We already have an agency running email. Does this conflict?
No. Most of this is infrastructure rather than campaign marketing — billing, reporting, the manual ops jobs — and it is usually the piece nobody at a campaign agency owns. Where the flow work genuinely overlaps, we hand your agency the audit and the sequence map and let them build against it.
We can’t actually tell you our profit by channel. Is that a problem?
It is the most common answer we get, and it is the first thing the audit fixes. A brand running eight to twelve marketing tools has eight to twelve versions of revenue, each attributed by the tool that sold it to you, and none of them carrying cost of goods or the real shipping charge. Until that is reconciled, every other decision on this page is being made on instinct.
Find out what you’re losing.
Before you commit to anything, we tell you exactly what you’re losing and what it costs to stop it. Two weeks. Fixed fee. Credited in full against any build you go ahead with.
- Fee
- $1,500–$3,000, fixed
- Duration
- Two weeks
- Credited
- In full, against any build
- You supply
- Read access + one 45-minute call