All segments

n8n Dashboard: Monitor Failed Runs Before Customers Do

An n8n dashboard should surface failed executions, queue depth and last-run timestamps, with a threshold that pages someone before a customer notices.

  • Published
  • Reading time 14 min read
  • Author Nafiul Hasan
n8n Dashboard: Monitor Failed Runs Before Customers Do. Diagram: what clears the floor. RUN n8n Dashboard: Monitor Failed RunsBefore Customers Do THE FLOOR pointerflow.com

Short answer

An n8n dashboard worth building surfaces failed executions, queue depth in queue mode, and the last successful run per workflow, pulled from n8n's own execution history and API rather than guessed at, with an alert that pages someone when a workflow has gone quiet for longer than its expected interval, not only when it errors loudly.

Why a green n8n execution list doesn’t mean the automation layer is healthy

An operator checks n8n once a day, sees mostly green in the execution list, and moves on. That habit works right up until the workflow that syncs abandoned-cart emails, or updates inventory across sales channels, or files a fulfilment request, stops running entirely, and nothing in that execution list tells you, because a workflow that never triggers leaves no failed execution behind. It leaves an absence, and an execution list built to show what happened doesn’t surface what didn’t.

An n8n dashboard earns its place on that argument: not a nicer view of the same execution history, but visibility into the automation layer’s own health, separate from whether any individual workflow’s last run succeeded. If you’re running Shopify Plus or a comparable subscription platform with enough workflows doing real operational work that a customer would notice if one stopped, the gap between “the execution list looks fine” and “every workflow that should have run today, ran” is where the expensive failures live. This article covers what to surface, where that data actually comes from, and where the alert threshold should sit so it pages someone before a customer emails support.

What you need before you build an n8n dashboard

You need read access to n8n’s own execution data, either through its REST API with an API key, or, for smaller setups, direct access to its underlying database. The API route is the one worth building against for anything you expect to keep running past a version upgrade or two, since n8n’s internal database schema has changed between releases and an API contract is more stable to build monitoring against than table structure you’re reading directly.

You need a place to send alerts that people actually check outside of n8n itself, whether that’s Slack, email, or a dedicated paging tool for anything urgent enough to wake someone up. This dashboard is only as useful as the alert path attached to it; a dashboard nobody looks at until something’s already broken is a record of the incident, not a way to catch it.

You need a list, even a rough one, of which workflows are customer-facing enough to justify a page versus which are internal-only and can wait for a morning check. Not every workflow deserves the same urgency, and building the alert threshold before deciding this tends to produce either alert fatigue from paging on everything, or silence on the one workflow that actually mattered.

If you’re running n8n in queue mode, with separate worker processes and a Redis-backed job queue, you also need access to that Redis instance or a queue-monitoring tool pointed at it, since queue depth isn’t something n8n’s own execution API surfaces on its own.

Decide what “healthy” means for your workflows

Before building anything, write down, per workflow, what a healthy state actually looks like. For a scheduled workflow, healthy means it ran, and ran successfully, within its expected interval; for a webhook-triggered workflow, healthy means it’s still receiving triggers at roughly the rate the source system sends them, which is a different check entirely from whether the runs it did receive succeeded. Conflating these two is the single most common gap in n8n monitoring setups: a dashboard built only to track execution success rate says nothing about a webhook subscription that’s silently expired and stopped delivering events at all.

Group workflows by what they actually do for the business, not by when they were built or which folder they live in. A workflow that updates a nightly reporting sheet and a workflow that triggers a fulfilment request when an order comes in both look identical in n8n’s execution list, one row per run, but a missed run on the first costs you a stale report; a missed run on the second costs a customer their order.

This grouping is what the alert threshold, covered later, actually keys off. Skipping it means every workflow gets treated the same, which in practice means either everything pages, and people start ignoring the pages, or nothing does, and the failures that matter get found by a customer instead of a dashboard.

Pull failed executions from n8n’s own data

n8n’s execution history, whether you’re on the default SQLite file or, for most production self-hosted setups, a PostgreSQL database you’ve pointed it at, is the primary record of what ran, when, and whether it succeeded. The stable way to pull this into a dashboard is n8n’s own REST API, authenticated with an API key generated from your instance’s settings, which exposes an executions endpoint you can filter by workflow and status.

Build a scheduled workflow, or a small external script if you’d rather keep monitoring logic outside n8n entirely, that queries this endpoint on a regular interval, say every five or ten minutes, and pulls executions with a failed or error status since the last check. For each one, record the workflow name, the failure time, and whatever error detail the API returns, into a datastore you control rather than relying on n8n’s own retention, since n8n’s execution data pruning settings clear old runs after a period you configure, and a dashboard tracking weekly or monthly trends needs history that outlives that pruning window.

Resist the urge to alert on every single failed execution the moment it happens. A workflow with a retry-on-fail setting that resolves on its second attempt doesn’t need a human involved; only a failure that exhausts its retries, or a workflow with no retry configured at all, is worth surfacing as an actionable failure rather than routine noise the retry mechanism already handled.

Track queue depth if you’re running queue mode

If your n8n instance runs in queue mode, with a Redis-backed job queue distributing executions across separate worker processes, queue depth, how many jobs are waiting for a worker to pick them up, is a health signal that failed-execution counts alone won’t give you. A queue depth that climbs steadily rather than draining back down means workflows are arriving faster than your workers can process them, and every workflow behind that backlog is running later than it should, even though none of them have technically failed yet.

Reading queue depth means querying the Redis instance n8n’s queue is backed by, either directly, since Bull and BullMQ, the queue libraries this kind of setup is typically built on, store queue state in specific Redis keys you can inspect, or through a generic queue-monitoring tool such as Bull Board, pointed at the same Redis instance, which gives a visual view without writing your own Redis queries. Either way, this data lives outside n8n’s own execution API, which is why teams running queue mode and monitoring only through n8n’s API miss this signal entirely.

Set a threshold for what counts as a problem based on your own worker throughput, not a number borrowed from another team’s setup, since the right depth depends on how many workers you’re running and how long your average execution takes. What matters more than the specific number is the trend: a queue depth that’s climbing over a sustained window, say the last thirty minutes, and shows no sign of draining is worth a look even if the absolute number seems small, because it means the backlog is compounding, not stable.

Record the last successful run per workflow

The last-successful-run check catches what failed-execution monitoring can’t: a workflow that’s stopped running entirely. Query n8n’s execution API per workflow, filtered to successful runs, and record the timestamp of the most recent one. Compare that timestamp against how often the workflow is supposed to run, whether that’s a fixed schedule or, for webhook-triggered workflows, a rough expected frequency based on how often the source system normally sends events.

A scheduled workflow set to run every fifteen minutes whose last successful run was two hours ago has a problem, whether or not its most recent attempt shows as a visible failure in the execution list. Common causes include a workflow that’s been deactivated by mistake, often during unrelated maintenance, an expired credential that’s blocking the trigger itself rather than a node inside the workflow, or, for webhook triggers, a subscription on the source system’s side that’s silently expired and stopped sending events at all.

Build this as its own check, separate from the failed-execution monitor, querying “when did each critical workflow last succeed” on a schedule, say every ten to fifteen minutes, and comparing that against each workflow’s expected interval. This is the closest thing to a direct answer to “is the automation layer actually working,” as opposed to “did the workflows that ran, run successfully,” which is a narrower and less useful question.

Build the alert that pages before a customer notices

The alert that matters most isn’t “a workflow failed.” It’s “a workflow hasn’t succeeded in longer than it should have,” because that catches both the loud failures and the silent ones, a stopped trigger, an expired credential, a deactivated workflow, that never generate a failed execution to alert on in the first place. This pattern, sometimes called a dead man’s switch, checks for the absence of an expected event rather than reacting only to events that did happen.

Set the threshold as a multiple of each workflow’s expected interval, not a fixed time for every workflow. A workflow that runs every fifteen minutes and a workflow that runs once a day need very different absolute thresholds even though the underlying question, has this workflow gone quiet for longer than it should have, is the same. As a starting, illustrative rule of thumb: for a workflow expected every fifteen minutes, three missed intervals, forty-five minutes of silence, is a reasonable point to page someone, rather than paging on the very first missed run, which risks false alarms from a single slow execution that would have finished on its own.

Tier the alert’s urgency by the workflow grouping you did earlier. A customer-facing workflow going silent past its threshold should page whoever’s on call immediately, through whatever tool already handles urgent operational alerts. An internal-only workflow going quiet can land as a Slack message for someone to check during business hours instead. Sending both through the same channel at the same urgency is how teams end up ignoring pages altogether, because the signal-to-noise ratio drops until nobody trusts the alert anymore.

Build the alert itself as its own small n8n workflow, or an external script if you’d rather keep it outside the system it’s monitoring, entirely, querying the last-successful-run data on a schedule and calling out through an HTTP Request node to whichever alerting tool you use. Keeping the monitor separate from the workflows it watches matters: if the monitoring workflow lives on the same n8n instance and that instance goes down entirely, nothing pages at all, which is worth at least being aware of even if a fully separate monitoring system isn’t proportionate to your scale yet.

Put it somewhere people actually look

A dashboard that only exists as a query someone has to remember to run doesn’t get checked until after something’s gone wrong. Build the actual visible dashboard, whether that’s a simple internal page, a spreadsheet that refreshes from the data you’re already logging, or a business intelligence tool you already use, somewhere it’s part of an existing daily routine, not a separate destination people have to remember exists.

Show failed executions by workflow over the last day and week, queue depth trend if relevant, and time-since-last-success per critical workflow, ranked so the workflow closest to breaching its threshold sits at the top. This ordering matters more than it sounds: a dashboard listing forty workflows alphabetically buries the one that’s actually close to a problem among thirty-nine that are fine.

Keep this dashboard separate from n8n’s own interface, even though the temptation is to just tell people to check n8n directly. n8n’s execution list is built for debugging a specific run, not for scanning the health of dozens of workflows at a glance, and asking people to use a debugging tool as a monitoring tool is why the habit of “check it once a day, see mostly green, move on” forms in the first place.

The step most teams get wrong

Most n8n monitoring setups, once a team builds one at all, correctly track failed executions. What they miss is the last-successful-run check, monitoring for absence rather than only for error. This gap is easy to miss precisely because it doesn’t produce any symptom until the day it matters: everything looks fine in a dashboard that only counts failures, because a workflow that’s stopped triggering entirely isn’t failing, it’s just not running, and a monitor built to catch failures has nothing to catch.

The fix isn’t complicated once you see the gap: a last-successful-run query, compared against each workflow’s expected interval, catches it. What makes it easy to skip is that it requires deciding, per workflow, what “expected interval” even means, which is more setup work than pulling a failed-execution count, and that extra decision is exactly the part teams defer and then forget to come back to.

Where does this monitoring data actually come from?

n8n’s execution data, failed and successful runs alike, lives in whichever database your instance is configured against, the default SQLite file for small or evaluation setups, or PostgreSQL for most production self-hosted instances handling real volume. The stable way to read it for a monitoring dashboard is n8n’s own REST API rather than querying that database directly, since the API contract is documented and versioned, while internal table structure has changed between n8n releases and isn’t guaranteed to stay stable.

Queue depth, when it applies, comes from Redis, not from n8n’s own API at all, since queue mode’s job queue is a separate piece of infrastructure n8n hands work off to. Reading it means either direct Redis queries or a queue-monitoring tool pointed at the same instance. Teams that build monitoring only against n8n’s API, without realising queue mode’s health data lives somewhere else entirely, end up with a dashboard that looks complete but is blind to a specific and common failure mode, a growing backlog that hasn’t yet produced a single failed execution.

What threshold should actually wake someone up?

The threshold that should page someone overnight is time-since-last-success on a customer-facing workflow exceeding a small multiple of its expected interval, not any single failed execution on its own. A single failed run that a retry mechanism resolves a minute later is routine; a customer-facing workflow that’s gone quiet for several missed intervals is not, and the difference between those two is exactly what a monitor built only around failed-execution counts can’t distinguish.

Queue depth deserves its own, separate threshold, keyed to sustained growth over a window rather than a single reading, since queue depth naturally fluctuates as jobs arrive and drain throughout the day. A depth that spikes and drains within a few minutes is normal load; a depth climbing steadily over thirty minutes or more with no sign of draining is a capacity problem worth waking someone up for, particularly if it’s affecting a customer-facing workflow’s timeliness.

Review these thresholds periodically rather than setting them once and forgetting them. As workflow volume grows, what counted as a normal queue depth six months ago may now be an early warning sign, and a threshold that never gets revisited either starts firing false alarms as normal volume grows past it, or stops catching real problems if volume patterns shift the other way.

How do you verify the dashboard itself is working?

Test it deliberately rather than trusting it because it looks complete. Deactivate a test workflow in a non-production n8n instance, or point a monitoring check at one with a known expected interval, and confirm the last-successful-run alert actually fires once that interval’s threshold passes. This is the one test that proves the dead-man’s-switch pattern works, as opposed to assuming a query that looks correct in isolation behaves correctly once it’s running on a schedule against real data.

Check that the alert reaches the right channel at the right urgency, not just that an alert fires somewhere. It’s common to build and test the underlying query, confirm it detects the problem correctly, and only discover during a real incident that the alert’s destination was misconfigured, or that it landed in a channel nobody was actually watching that week.

Periodically audit the workflow list your monitoring covers against the actual list of workflows running in n8n. A new workflow built without being added to the monitoring dashboard’s tracked list is invisible to it entirely, however good the underlying monitoring logic is, and this gap tends to open quietly as a team adds workflows faster than anyone remembers to register them for monitoring.

Getting this right is a reporting and analytics problem as much as an n8n problem: the same discipline of deciding what to surface, where the data actually lives, and what threshold justifies waking someone up applies to any operational system a store depends on, which is the kind of visibility work Pointerflow’s reporting and analytics covers for stores running their automation layer in-house.

Sources

  • No external figures are quoted in this article. It’s written from n8n’s documented features (the execution list, REST API, Error Workflow, queue mode and its Redis-backed job queue, and the optional Prometheus-compatible metrics endpoint) described at the level of mechanism, since specific limits, retention windows and environment variable names vary by n8n version and are best checked against current documentation for the version in use.

Frequently asked

What data points does an n8n dashboard need at minimum?

A useful n8n dashboard shows failed executions by workflow, queue depth if you're running queue mode, and the time since each critical workflow's last successful run. Execution counts and success rates matter less than these three, because a workflow that's stopped triggering entirely won't show up as a failure at all, only as silence.

Does n8n have a built-in dashboard for this?

n8n's own interface shows an execution list with status per run, filterable by workflow, which covers ad hoc debugging well. It doesn't aggregate failure counts, queue depth or per-workflow last-run timestamps into one view by default, which is the gap a dedicated monitoring dashboard, built from n8n's API or metrics endpoint, fills.

Does querying n8n's API for monitoring slow the instance down?

A monitoring workflow polling n8n's REST API every five or ten minutes adds negligible load next to the workflows n8n already processes. Keep polling to a reasonable interval rather than querying every few seconds, since the signals that matter, failures and last-run timestamps, don't change fast enough to justify tighter polling.

What does queue depth actually measure day to day?

Queue depth applies when n8n runs in queue mode, where a Redis-backed job queue distributes executions across worker processes. Queue depth is how many jobs are waiting for a worker to pick them up. A depth that keeps climbing rather than draining means workflows are queuing faster than your workers can process them.

Is queue-mode monitoring necessary for a single-instance n8n setup?

Single-instance n8n installs (no separate worker processes, no Redis queue) don't have a queue depth to monitor, but they still have failed executions and per-workflow last-run timestamps, which cover most of what matters. The main added risk without queue mode is that one long-running or stuck execution can block others behind it.

What's a dead man's switch in workflow monitoring?

A dead man's switch is a separate check that alerts when an expected event hasn't happened, rather than reacting to an event that did. Applied to n8n, it means alerting when a workflow that's supposed to run every 15 minutes hasn't logged a successful execution in longer than that, which catches a workflow that's stopped triggering entirely.

Why is a silent workflow failure worse than a loud one?

A loud failure, an execution that errors and shows red in n8n, is visible the next time anyone checks. A silent failure, a trigger that's stopped firing, a webhook subscription that's expired, a scheduled workflow that's been deactivated by mistake, produces no error at all, just an absence, and absences don't show up in a dashboard built only to count failures.

Should every workflow page someone on failure?

No. A low-stakes internal reporting workflow failing overnight can wait for a morning review, while a workflow that syncs order data to a fulfilment system needs a page within minutes. Tier your workflows by customer impact and set alert urgency per tier, rather than treating every red execution the same way.

Can n8n's error workflow feature replace external monitoring?

An Error Workflow catches a run that fails while executing, which is useful, but it can't catch a workflow that never started, since there's no execution to trigger it from. External monitoring, checking for the absence of expected executions rather than only reacting to failed ones, covers the gap Error Workflow alone leaves open.

What tool should send the actual alert?

Most teams route n8n monitoring alerts through whatever's already used for other operational paging, Slack for low-urgency notices, a dedicated paging tool for anything that should wake someone up overnight. n8n can call either through an HTTP Request node in the monitoring workflow itself; the choice depends on your existing on-call setup, not on n8n.

How much history should you keep for n8n executions?

n8n's own execution data pruning settings control how long runs stay queryable through its interface and API before being cleared, and the default retention is short enough that a monitoring dashboard checking daily or weekly trends needs its own separate log rather than relying on n8n to keep history indefinitely. Check your instance's current pruning configuration.

Does n8n expose Prometheus metrics?

n8n ships an optional Prometheus-compatible metrics endpoint you can enable through environment configuration, useful if you already run Prometheus and Grafana for other infrastructure. The exact environment variable and metric names have changed across n8n versions, so check the documentation for your installed version rather than an older guide.

Next step

Is this your reporting & analytics problem, or a symptom of another one?

Bring your numbers — the churn split, the decline rate, whatever your flows are earning — and we will tell you which of them is the expensive one.

Book a call →