
How to Run n8n in Production Without Breaking Workflows (2026)
A practical guide to running n8n reliably: queue mode with Redis and Postgres, error workflows that alert you, duplicate-safe webhooks, secrets, AI agent guardrails and backups.
Most n8n setups start on a small server with default settings, and for a while that's fine. Then a webhook gets dropped during a restart, the CRM ends up with three copies of the same lead, or an AI step hangs a workflow for an hour. The question changes from "does it work?" to "will it keep working when nobody is watching?" This is the setup we use when an n8n instance moves from experiment to something the business depends on.
How do you run n8n reliably in production?
Run n8n on PostgreSQL rather than SQLite, and switch to queue mode (a main instance plus workers, with Redis as the queue) once load or isolation demands it. Give every production workflow an error workflow that alerts a person, make webhooks ignore duplicate deliveries, set and safely store your encryption key, put guardrails on AI Agent steps, and monitor health, back up the database and keep workflows in version control.
Step 1: choose the right hosting setup

n8n has two execution modes. Regular mode, the default, runs the editor, triggers, webhooks and every execution in one process. For a handful of low-volume workflows that is perfectly adequate.
Queue mode splits the work. The main instance handles the editor, timers and incoming webhooks, creates each execution and puts it on a queue in Redis. Separate worker processes pick jobs off the queue, run them and write the results back to the database. We move a project to queue mode when:
Heavy executions (large files, long loops, slow AI calls) start slowing the editor or delaying other workflows.
Webhook traffic arrives in bursts and adding workers is cheaper than buying a bigger server.
You want to update workers without taking the editor and webhook endpoints offline.
A few rules apply either way:
Use PostgreSQL. n8n's docs advise against queue mode on SQLite, and Postgres gives you standard backup tools and an easy path to scaling later.
Plan for files. Queue mode doesn't support binary data stored on the local disk, so workflows that keep files need n8n's S3-compatible external storage.
Run code in separate task runners. Code nodes should run in external task runners, one sidecar per n8n process, on the same version as n8n itself. Without that isolation, anyone who can edit a workflow could potentially read your credentials and environment.
Pin the version. Don't run "latest" in production. n8n ships releases frequently, so upgrade deliberately, after a backup.
Keep the encryption key identical everywhere. The main instance and every worker need the same N8N_ENCRYPTION_KEY, or workers can't read stored credentials.
Worker concurrency defaults to 10 jobs per worker, and n8n recommends keeping it at 5 or more, because lots of low-concurrency workers can exhaust the database connection pool. For very heavy webhook traffic, n8n also supports dedicated webhook processes behind a load balancer.
Step 2: set up error handling that tells you something broke

The most damaging n8n failures are the silent ones. n8n gives you three layers of error handling, and we use all three.
Retry on fail. Each node's settings include Retry On Fail, Max Tries and Wait Between Tries. Turn it on for nodes that call external APIs, where most failures are timeouts, rate limits or brief outages. For rate limits, wait longer than the provider's limit window.
A per-node On Error setting. The choices are Stop Workflow (the default), Continue, and Continue (using error output). Plain Continue is rarely what you want, because the next node runs as if nothing happened and the failure disappears. The error output option sends failed items down a separate branch, so one bad record can go to a review list while the rest of the batch carries on.
An error workflow. Build one workflow that starts with the Error Trigger node and set it as the Error Workflow in the settings of every production workflow. When something fails, it receives the workflow name, the node that failed, the error message and a link to the execution, and can post all of that to Slack or email.
A few details catch people out. The error workflow only fires on automatic runs, so you can't test it with a manual execution. The link to the failed execution only exists if failed executions are saved, so keep that setting on. And when an API returns a "success" response with the wrong data, the Stop and Error node lets you turn that into a real failure that triggers the alert.
Step 3: make webhooks safe to receive twice

Webhook senders retry. Payment providers, form tools and CRMs will deliver the same event again if your endpoint was slow or a deploy landed at the wrong moment, and every retry can become a duplicate order, lead or invoice. The fix is to make each webhook workflow idempotent: processing the same event twice has the same result as processing it once.
Before doing any work, check the unique ID the sender includes with each event (or a fingerprint of the payload) against the IDs you've already handled. n8n's Remove Duplicates node can do this across executions and keeps 10,000 items of history by default. For money or customer records, we prefer a small database table we control, where the event ID is the primary key: if the insert succeeds the workflow continues, and if the ID already exists it stops. Because the database enforces uniqueness, two workers can't process the same event at the same time.
This matters most in order and stock workflows. Our guide to syncing Shopify orders and inventory with n8n shows where duplicates tend to creep in.
Step 4: protect your secrets and separate environments

The encryption key is the one thing you can't lose. n8n encrypts every saved credential with it. If you don't set it yourself, n8n generates one and stores it in its data folder, which breaks as soon as you add a worker, rebuild a volume or restore onto a new server. Set it explicitly and keep a copy in a password manager or secrets vault. A database backup without its key gives you credentials nobody can decrypt.
Never paste API keys into nodes or expressions. They leak into workflow exports, git history and screenshots. Use n8n's credentials, which are encrypted at rest.
Keep infrastructure secrets out of plain environment variables. Most n8n settings can read their value from a file instead (add _FILE to the setting name), which works with Docker and Kubernetes secrets.
Block environment access from workflows. Setting N8N_BLOCK_ENV_ACCESS_IN_NODE to true stops Code nodes and expressions from reading environment variables, which may include your database password.
Separate dev, staging and production completely. Each gets its own instance, database, encryption key and credentials. n8n's source control feature (Business and Enterprise plans) links instances to git branches, and Enterprise can connect each instance to an external secrets vault. On the community edition, export and import workflows between environments and create credentials separately in each.
Step 5: put guardrails on AI Agent steps

The AI Agent node connects a language model to tools and lets the model decide which tools to call. That means part of your workflow's behaviour is chosen by the model, so treat its output as untrusted until it has been checked.
Constrain the output. Turn on Require Specific Output Format and attach a Structured Output Parser with a JSON schema, then validate again after the agent: allowed values, IDs that exist, sensible amounts. Anything that fails goes to Stop and Error, so your error workflow hears about it.
Scope tools narrowly. A read-only CRM lookup is low risk. A tool that updates or deletes records is not. Give agent tools their own limited-permission credentials where the system allows it.
Put a human in front of writes. n8n lets you add a human review step in front of selected tools, with approval requests sent to Slack, email, Teams or other channels. The workflow pauses until someone approves or denies the action. We start new agents with review on every write and relax it only when the logs justify it.
Cap time and cost. Keep Max Iterations (10 by default) so a confused model can't loop through tool calls, and set a workflow timeout or an instance-wide default with EXECUTIONS_TIMEOUT. The default wait for AI providers is an hour, far too long for anything a customer is waiting on.
Log prompts and outputs. Execution data gets pruned, so write the input, prompt version, output and tool calls to your own table, without personal data you don't need. When someone asks why a lead was marked as spam, you'll have the answer.
Step 6: monitor, back up and version your workflows

Task | What to do |
|---|---|
Health checks | Monitor the main instance's readiness endpoint, which confirms the database is connected. In queue mode, enable worker health checks too. Keep the optional metrics endpoint off the public internet. |
Execution pruning | Pruning is on by default (14 days or 10,000 executions). Busy instances should set both limits deliberately so the database doesn't balloon. |
Database backups | Back up Postgres nightly with standard tools and copy the files off the server. Include n8n's data folder, which holds the encryption key. |
Restore tests | Restore onto a spare server every so often. An untested backup is only a hope. |
Workflows in git | Export workflows on a schedule with n8n's command-line tool and commit them, so you can see what changed and roll back. Never commit decrypted credentials. |
n8n recommends a full backup before every upgrade. Also note that imported workflows arrive deactivated by default, so plan how they get switched on when you move them between environments.
Common mistakes
Running production on SQLite and the default settings it was tested with.
Letting n8n generate the encryption key, then losing it in a server rebuild.
Using "Continue" on errors, so failures vanish instead of raising an alert.
Processing every webhook delivery as new, which turns retries into duplicate records.
Giving an AI agent write access to the CRM with no validation or human review.
Backing up the database but never testing a restore.
Want help running n8n in production?
Ryven builds and maintains n8n automations for companies in the US, UAE, UK and Australia, from first workflows to queue-mode setups with monitoring and AI steps. If you're comparing partners, read how to choose an n8n automation agency, see our n8n automation services, or book a free consultation and we'll review your current setup.



