What are Klaviyo webhooks, and when do you actually need one?
Klaviyo webhooks move data in and out of Klaviyo the moment something happens, rather than waiting for the next scheduled sync. Inside a flow, a webhook action posts a JSON payload to a URL you control the instant a profile reaches that step, carrying identifiers and event context you can act on immediately in a warehouse system, a helpdesk, or a scoring model you run yourself. That is the sending side. The receiving side, getting data from somewhere else into Klaviyo through a webhook, is a different mechanism built around an endpoint you or a middleware tool operate, not a raw inbound URL Klaviyo hands you.
You need a webhook only where timing genuinely matters and no native integration already covers the gap. If a tool already has a listed Klaviyo integration, or the deadline is same-day rather than same-second, a scheduled export or a batch call through the API does the job with far less to maintain. A webhook is infrastructure you now own: retries to handle, an endpoint to keep online, an auth token to rotate. Build one because nothing else moves fast enough for the use case, not because a real-time call sounds more capable than a nightly export. Teams doing this well at $3M–$30M in revenue tend to have exactly one or two webhook pairs in production, each solving a problem an export genuinely cannot, not a webhook wired to every flow because the option exists.
How do you set up an outbound webhook inside a Klaviyo flow?
Add the webhook action at the exact point in the flow where the call should fire, not earlier in the sequence where a profile might still exit through a flow filter. Set the destination URL, choose the payload shape, and add whatever headers the receiving system expects, typically an authentication token and a content-type declaration. Save the flow in draft and send at least one test profile through it against a sandbox endpoint that simply logs whatever arrives, so you can see the real payload before a live customer’s data goes anywhere.
Order matters more than it looks. A webhook action placed before a flow filter fires for every profile that enters the flow, including ones the filter would have excluded a few steps later; a webhook placed after the filter only fires for profiles that actually qualify. Time-delay steps before the webhook change the payload’s timestamp relative to the triggering event, which matters if the receiving system uses that timestamp to sequence records. Once the action is live, changing the destination URL or the payload template is a production change to a system another team may depend on, so treat it with the same care as an API endpoint change, not as a flow tweak.
What fields does a Klaviyo webhook payload carry?
A webhook payload typically carries the profile’s identifiers, the event or flow context that triggered the call, and a timestamp, shaped by whichever payload template you set on the action itself. Beyond that base shape, the exact fields available to include depend on your account, the flow, and the data attached to the profile at that point, so the payload builder inside your own flow is the source of truth, not a fixed schema copied from another integration.
Two things go wrong here more than anything else. First, a payload template built by one person and later edited by another drifts quietly: a field gets renamed, a property gets removed because it looked unused, and the receiving system starts getting nulls where it expected values, with no error thrown on either side because the call still succeeds. Second, nested or array-shaped data in the payload gets flattened incorrectly on the way into a downstream system that expects a flat record, silently dropping everything past the first item in a list. Log the raw payload on the receiving end for at least the first few weeks after launch, so a schema change shows up as a diff you can see rather than a support ticket three months later asking why a field has been empty since some point nobody can pin down.
How do you get data from another system into Klaviyo through a webhook?
You receive the inbound webhook on your own endpoint or a middleware platform, then translate it into a server-side call into Klaviyo, rather than pointing another system’s webhook straight at Klaviyo and expecting it to understand an arbitrary payload. This distinction catches people who assume “Klaviyo supports webhooks” means it exposes a generic inbound URL that accepts any JSON shape and maps it to a profile automatically. It does not work that way; something in the middle has to know both schemas.
That middle layer can be code you own, a low-code platform, or a purpose-built connector, and the right choice depends on volume and how much transformation the data needs before it is useful in Klaviyo. High-volume, simple pass-through data, an order placed event from a platform with a clean webhook, suits a small piece of code you maintain. Lower-volume data from a system with a messier webhook payload, a helpdesk closing a ticket with several nested fields, often suits a middleware tool where the mapping is visual and a non-engineer can adjust it later without a deploy. Either way, confirm the current recommended pattern and any field-mapping constraints in Klaviyo’s own integration documentation before building against assumptions, because the specifics of how events and profile properties are accepted are the kind of detail that changes between account tiers and API versions.
What happens when Klaviyo can’t deliver a webhook on the first try?
Klaviyo retries a failed delivery a limited number of times over a widening interval before marking it failed and stopping. That is the retry-ladder shape: an immediate attempt, then a second attempt after a short wait, then further attempts spaced further apart, until the last attempt either succeeds or the delivery is given up on. The exact retry count and the exact wait between attempts are not values worth quoting from memory here; they are the kind of operational detail that can change between plans or product updates, so check Klaviyo’s current documentation for the retry window that applies to your account before you design alerting around a specific number.
What you control is what happens on your side while the ladder runs. If your endpoint is briefly overloaded and returns a timeout on the first attempt, a retry a short time later often succeeds without you doing anything, which is the entire point of the ladder existing. The risk is treating every failed first attempt as an emergency, when the real emergency is only the final failed attempt, the one Klaviyo does not retry again. Build your monitoring around that last state, not every individual failed call.
How do you build idempotency into the endpoint that receives Klaviyo’s webhooks?
Idempotency means a retried delivery produces the same result as the original call, not a second copy of it. Give every incoming event a stable identifier, built from the profile, the flow action and a timestamp if nothing more specific is available, and check that identifier against what you have already processed before writing anything downstream. A retried delivery should hit the same identifier, get recognised as already handled, and get discarded quietly rather than creating a duplicate order note, a duplicate support ticket, or a duplicate row in whatever system is on the receiving end.
Store the identifiers somewhere durable and queryable, not in memory, because a retry can arrive minutes after the original call, well past the life of any in-process cache. A small table with the identifier, a received timestamp, and a processing status is enough for most volumes; at higher volumes, an expiring key-value store keyed on the identifier does the same job faster. Decide how long to keep those identifiers around, too. Keep them forever and the table grows without bound; expire them too soon and a very late retry, one that arrives after Klaviyo’s own retry window but is replayed manually during an incident, slips past the check and creates the duplicate you built the whole system to prevent.
What should your endpoint send back, and how quickly?
Return a success status as soon as you have safely queued the payload, not after you have finished acting on it. If your endpoint does anything slow, writing to a warehouse, calling a third API, updating an order management system, do that work asynchronously after you have already told Klaviyo the delivery succeeded. A slow synchronous response that times out looks exactly like a failure to Klaviyo, even if your system eventually finished the work correctly, and that mismatch between “we got it” and “the call took too long to answer” is one of the most common causes of duplicate processing once retries are in play.
The pattern that holds up at volume is a thin receiving layer and a separate worker. The receiving layer does three things and nothing else: verify the request came from where it claims to, write the raw payload to a queue or a log with its idempotency identifier, and return a success response immediately. A worker process, running separately, reads from that queue and does the actual work, on its own timeline, with its own retry logic if something downstream is unavailable. This separation means a slow downstream system never causes Klaviyo to see a timeout, because the two problems, “did we receive it” and “did we finish processing it,” stop being the same problem.
Which three failure modes actually break a Klaviyo webhook integration?
Payload-template drift causes the first failure. Someone edits the flow’s webhook action, renames or removes a field to tidy it up, and the change ships without anyone checking what the receiving system expects. The call still returns a success status, because the endpoint accepted the request; the failure only shows up downstream, as a null where a value used to be, and it usually surfaces as a data-quality complaint weeks later rather than an alert on the day it happened.
Receiver timeouts causing duplicates instead of clean retries are the second failure. An endpoint that does slow synchronous work before responding times out under load, Klaviyo’s retry fires, and if the endpoint is not idempotent the retry finishes the original slow work a second time, on top of whatever partial work the first, timed-out attempt already completed. This is the failure mode that produces two order notes, two loyalty-point credits, or two support tickets for one event, and it gets worse under load rather than better, because load is exactly when timeouts and retries both spike together.
Auth or certificate rotation on the receiving endpoint causes the third failure. A token used in the webhook’s headers gets rotated as part of routine security hygiene, or a TLS certificate on the receiving server expires and renews with a new fingerprint, and nobody updates the flow action or the receiving service’s allowlist at the same moment. Every delivery fails from that point, Klaviyo works through its retry ladder and gives up, and because there is no default alert tied to a webhook action failing, the gap can run for hours or days before someone notices data has stopped arriving. This is the failure mode worth building a specific check for, because it is invisible everywhere except in a direct comparison between what should have arrived and what actually did.
What do you do once you’ve confirmed a webhook delivery went missing?
Confirm the gap first, then size it, then decide whether to backfill. Compare a count from the sending side, profiles that reached the webhook step in flow analytics for the affected window, against a count on the receiving side, rows written in the same window. If flow analytics shows more profiles reaching that step than your system recorded, you have a real gap and a rough size for it, not just a suspicion.
Backfilling depends entirely on what data is still available. If the missed event is still present somewhere in Klaviyo, in flow analytics, in a segment, or in exportable profile activity, you can often reconstruct the missing payloads by exporting the relevant profiles and events through the API rather than the webhook, since the API and the webhook are two doors into overlapping data rather than the same door twice. If the underlying event no longer exists anywhere queryable, the honest answer is that it is gone, and the fix is process rather than recovery: shorten the window between a delivery failing and someone noticing, so the next gap is caught in hours rather than weeks. A calculator like the one at /tools/flow-revenue-calculator is a reasonable way to put a rough number on what a gap in a revenue-driving flow is worth while you are deciding how much engineering time the backfill deserves.
How do you monitor a webhook pipeline so a gap doesn’t sit unnoticed for a week?
Monitor the comparison, not the individual call. A dashboard that just shows “webhook calls received: yes” tells you the endpoint is up; it does not tell you whether the count matches what Klaviyo actually sent. Build a daily or hourly job that pulls a count from flow analytics for the step in question and a count from your own received-and-processed log for the same window, and alert when the two diverge by more than a small tolerance. This is the single check that catches all three failure modes above, because each of them produces the same symptom: fewer records on the receiving side than the sending side expected.
Log enough on the receiving side to answer a question you have not thought of yet. At minimum, keep the raw payload, the idempotency identifier, the timestamp received, and the outcome of processing, for a rolling window long enough to cover at least one full billing or reconciliation cycle. When someone asks three weeks from now why a customer’s loyalty points look wrong, that log is the difference between answering in five minutes and re-running a data audit across every system involved. Treat a webhook endpoint that has been silently returning success while dropping data as the worst version of this failure, because uptime monitoring alone will never catch it; only the comparison against the sending side will.
When should data move through a webhook instead of the Klaviyo API?
Use a webhook when the trigger is an event inside a flow and the receiving system needs to react within seconds, not minutes, of that event happening; that is the case a webhook exists to solve and an API poll cannot match without hammering the API on a tight schedule. Use the API instead for anything that can tolerate a delay, for bulk operations that would mean firing hundreds of individual webhook calls in a burst, and for reconciliation, backfills and audits, where you want to pull a known, bounded dataset rather than depend on every individual call having arrived correctly over time.
Brands scaling past a handful of flows, the kind this article assumes as its reader, at $3M–$30M in revenue on Shopify Plus or a comparable paid platform, tend to end up running both at once: a small number of webhooks carrying the handful of truly time-sensitive events, and scheduled API calls carrying everything else, including the reconciliation job that checks the webhooks are still working. A brand still running a handful of manual flows and no real integration layer is not the audience for this decision yet; get the flows themselves right, covered in more depth at /for/scaling-brands, before adding a real-time integration layer on top of them.
Webhook reliability is ultimately a lifecycle flows problem, not an integrations problem on its own: every payload that drifts, times out, or silently stops arriving is a flow that stops doing what the business assumes it is still doing. That is the territory covered at /services/lifecycle-flows, where the flows themselves, not just the pipes connecting them to other systems, get the ongoing attention a webhook integration alone cannot give them.
Sources
- Klaviyo, 41% of email revenue from automated flows across 183,000+ brands (vendor-reported), cited to support the framing that flow reliability, including webhook-connected flows, carries a meaningful share of email revenue.