Self-hosting SilenceWatch
{/* Generated from docs/self-hosting.md by scripts/prepare.mjs. Edit that file. */}
Everything here works on one small VPS. The whole system is a Node process and a PostgreSQL database; there is nothing else to install, scale or watch.
Requirements
Section titled “Requirements”- Docker with Compose v2
- 1 vCPU and 1 GB of RAM handles thousands of checks comfortably
- PostgreSQL 14 or later (16 recommended). The Compose file brings its own
Install
Section titled “Install”git clone https://github.com/liliang-dev/SilenceWatch.gitcd SilenceWatchcp .env.example .envFill in at least:
SECRET_KEY=$(openssl rand -hex 32) # all signing keys derive from thisPOSTGRES_PASSWORD=$(openssl rand -hex 24)BASE_URL=https://watch.example.com # what your users actually typeThen:
docker compose up -dCreate the first account at BASE_URL. On an empty instance sign-up is always
allowed, whatever SIGNUP_ENABLED says — otherwise a fresh install could never be
bootstrapped. Once your team is in, set SIGNUP_ENABLED=false and restart.
Configuration
Section titled “Configuration”Every setting is an environment variable. The server validates them at startup and refuses to boot on an unsafe or incoherent configuration rather than running in a degraded state you would discover during an incident.
The minimum
Section titled “The minimum”Three values. With any of them missing the instance does not start, and says which one:
| Variable | Set it to |
|---|---|
SECRET_KEY |
openssl rand -hex 32. Every signing key derives from it, so rotating it signs everyone out |
POSTGRES_PASSWORD |
openssl rand -hex 24. Compose builds DATABASE_URL from it |
SMTP_URL |
smtp://user:password@relay.example.com:587. Or a POSTMARK_TOKEN / BREVO_API_KEY with the matching EMAIL_PROVIDER |
SMTP_URL is in that list because docker-compose.yml defaults
EMAIL_PROVIDER to smtp, and a mail provider with no transport is a
configuration that would accept alerts and drop them. Nothing else on this page
has to be set.
One more you should set anyway, whose default is only right on a laptop:
| Variable | Default | Why it matters |
|---|---|---|
BASE_URL |
http://localhost:8080 |
Ping URLs and alert links are built from it, so a wrong value hands out links nobody can follow |
Which makes the smallest genuinely usable .env four lines:
SECRET_KEY=…POSTGRES_PASSWORD=…SMTP_URL=smtp://user:password@relay.example.com:587BASE_URL=https://watch.example.comDATABASE_URL is deliberately not in that list. The server does require it,
but docker-compose.yml and docker-stack.yml both assemble it from
POSTGRES_PASSWORD — you only set it yourself when pointing at a PostgreSQL the
Compose file did not start.
EMAIL_PROVIDER=console is not usable in production. It prints alerts to
the log instead of sending them, and the published image sets
NODE_ENV=production, where the server refuses to boot with it.
Alerting
Section titled “Alerting”| Variable | Default | Meaning |
|---|---|---|
EMAIL_PROVIDER |
console |
console, smtp, postmark, brevo |
SMTP_URL |
— | smtp://user:pass@host:587 (STARTTLS required) or smtps://…:465 |
POSTMARK_TOKEN / BREVO_API_KEY |
— | API token for the matching provider |
EMAIL_FROM, EMAIL_FROM_NAME |
— | Sender identity |
NOTIFICATION_MAX_ATTEMPTS |
6 |
Retries before a delivery is abandoned |
NOTIFICATION_TIMEOUT_MS |
10000 |
Timeout for each outbound alert |
ALLOW_PRIVATE_NOTIFICATION_TARGETS |
false |
Allow alerts to reach private addresses |
console prints alerts to the log instead of sending them, which is useful in
development and unacceptable in production — the server refuses to start with it
when NODE_ENV=production.
Do not self-host SMTP for alerts. Deliverability is a reputation game you have no reason to play; the day an alert matters is the day you do not want it in a spam folder. Relay through a provider or through a relay you already trust.
Sign-up
Section titled “Sign-up”| Variable | Default | Meaning |
|---|---|---|
SIGNUP_ENABLED |
true |
When false, only the first account can be created |
EMAIL_VERIFICATION_REQUIRED |
false |
Require a proven address before sign-in |
EMAIL_VERIFICATION_TTL_HOURS |
24 |
How long a confirmation link stays valid |
UNVERIFIED_ACCOUNT_TTL_DAYS |
7 |
Unconfirmed accounts are deleted after this; 0 keeps them |
SIGNUP_POW_DIFFICULTY |
0 |
Proof-of-work bits required to register; 0 disables it |
SIGNUP_POW_TTL_SECONDS |
600 |
Lifetime of an issued challenge |
SIGNUP_BLOCK_DISPOSABLE_EMAIL |
false |
Reject known throwaway mailbox domains |
SIGNUP_BLOCKED_EMAIL_DOMAINS |
— | Extra domains to reject, comma-separated |
SIGNUP_MAX_PER_NETWORK_PER_HOUR |
0 |
Accounts per hour per network prefix; 0 disables it |
Everything below the first line is off by default and stays that way for most
self-hosters. If your instance is on a private network, or you set
SIGNUP_ENABLED=false once the team is in, you have already solved the problem
these settings address — none of them is worth turning on.
They exist for instances whose sign-up form is reachable by strangers.
abuse-prevention.md explains what each one actually
buys, with measurements, including the ones that are not worth what people
expect.
Turning EMAIL_VERIFICATION_REQUIRED on also makes registration
enumeration-safe: the API stops revealing whether an address already has an
account. The server refuses to boot with it enabled and EMAIL_PROVIDER=console
— nobody could ever confirm an address, so every new account would be locked out
of an instance that otherwise looks healthy.
Plans and quotas
Section titled “Plans and quotas”A self-hosted SilenceWatch has no limits, and nothing here needs setting.
QUOTAS_ENABLED is off, user.plan stays null, and null means unlimited on
every axis. This section exists because the hosted deployment runs the same
code; it is documented so you can see that it does, and so you can use it if you
run SilenceWatch for other people.
| Variable | Default | Meaning |
|---|---|---|
QUOTAS_ENABLED |
false |
Master switch. Off means unlimited, always |
DEFAULT_PLAN |
free |
Plan given to a new account when quotas are on |
PLAN_LIMITS |
{} |
JSON: what each plan name is allowed |
QUOTA_RECONCILE_INTERVAL_MS |
300000 |
How often accounts are matched against their plan |
QUOTAS_ENABLED=truePLAN_LIMITS='{ "free": {"checks": 10, "projects": 3, "channelsPerProject": 3, "retentionDays": 7}, "supporter": {"checks": 10, "projects": 3, "channelsPerProject": 3, "retentionDays": 30}, "pro": {"checks": 100, "projects": 20, "retentionDays": 90}, "max": {"checks": 1000}}'An omitted key is unlimited: max above has no project ceiling and no retention
cap. An unknown plan name is also unlimited — a typo in this JSON, or a plan
renamed on the billing side, should briefly give someone too much rather than
lock a paying customer out of their own monitoring.
Which plan an account is on is the plan column on user. This repository never
writes it and knows nothing about prices, payment or subscriptions; whatever does
your billing sets the column, and the reconciler picks the change up within
QUOTA_RECONCILE_INTERVAL_MS.
Checks are counted across every project the account owns, so a second project does not reset the allowance. On a downgrade the excess checks are paused — newest first, deliberately paused checks untouched, and the account emailed the list — and they resume on their own when the account moves back under its limit.
Security and recovery
Section titled “Security and recovery”| Variable | Default | Meaning |
|---|---|---|
PASSWORD_RESET_TTL_MINUTES |
60 |
Lifetime of a reset link |
AUDIT_RETENTION_DAYS |
365 |
How long the audit trail is kept |
The audit trail records sign-ins and failed sign-ins, password changes and resets, API key and alert channel changes, check deletions, ping-key rotations and quota pauses. It is readable in Settings → Security activity by project admins, and never contains a token, a key or a channel’s configuration.
A leaked ping URL no longer means recreating the check: Check → ⋯ → Rotate ping URL issues a new one and keeps the history. Every job still calling the old URL will be reported as down, which is the point — update them first.
Detection and retention
Section titled “Detection and retention”| Variable | Default | Meaning |
|---|---|---|
DETECTION_INTERVAL_MS |
10000 |
How often the detection loop runs |
DETECTION_BATCH_SIZE |
200 |
Checks claimed per batch |
PING_RETENTION_DAYS |
90 |
Ping history kept (per-project override available) |
PURGE_CRON |
17 3 * * * |
When the purge runs, in UTC |
PING_RATE_LIMIT_PER_MINUTE |
120 |
Per ping key, so a looping job cannot saturate the database |
PING_BODY_MAX_BYTES |
10000 |
Ping bodies are truncated to this |
Behind a reverse proxy
Section titled “Behind a reverse proxy”Set TRUST_PROXY only when a proxy really is in front, and set it to the
proxy’s address or CIDR rather than true:
TRUST_PROXY=10.0.0.0/8A bare true trusts the hop count from anything that can reach the server, so
anyone bypassing the proxy can claim any client address they like — and every
per-address control (rate limits, sign-up velocity, the audit trail) is only as
good as that value. Production refuses to boot on TRUST_PROXY=true for
exactly this reason.
Getting it wrong in either direction is otherwise silent, so the server samples its first requests and says so in the log: trusting a proxy whose header never arrives, or refusing to trust one whose header always does. The second is the quieter failure — it throttles the entire internet as if it were one visitor.
A minimal nginx front:
location / { proxy_pass http://127.0.0.1:8080; proxy_set_header Host $host; proxy_set_header X-Real-IP $remote_addr; proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; proxy_set_header X-Forwarded-Proto $scheme;}docker compose up -d serves plain HTTP on PORT and nothing else. That is
deliberate: plenty of instances sit behind an nginx, a Traefik or a company load
balancer that already terminates TLS, and a certificate this side of it would be
one more thing to renew for no gain.
If nothing is in front, Compose can bring its own:
docker compose --profile tls up -dThat adds Caddy, which obtains a certificate and renews it without being asked.
The profile is why the plain path is unaffected — without --profile tls the
service does not exist, and docker compose up -d starts exactly what it always
did.
It needs two things that are not Docker’s to give: SILENCEWATCH_DOMAIN must
resolve to this machine, and ports 80 and 443 must be reachable from the
internet. Both are how Caddy proves the name is yours; a firewall or a DNS
record still pointing at a parking page fails certificate issuance, not the
container.
Four settings move together, and three of the four are easy to forget:
SILENCEWATCH_DOMAIN=status.example.com # the name on the certificateBASE_URL=https://status.example.com # or every ping URL you hand out is wrongBIND_ADDRESS=127.0.0.1 # stop publishing 8080 to the worldTRUST_PROXY=172.28.0.0/16 # the Compose network Caddy speaks fromBIND_ADDRESS matters more than it looks. Left at the default the application’s
own port stays open beside the certificate — a plain-HTTP way around it, on
which requests arrive with no X-Forwarded-For at all.
TRUST_PROXY takes the network rather than true, for the reason above: in
production a bare true is refused. The Compose network is given a fixed subnet
precisely so it can be named here; change it with DOCKER_SUBNET if
172.28.0.0/16 collides with something on your host.
Under Swarm there is no profile — Caddy is part of docker-stack.yml and always
deployed, since a swarm reached over the internet wants TLS anyway. The subnet
variable is SWARM_SUBNET, defaulting to 10.20.0.0/16.
A second name for a static site (the hosted service’s layout)
Section titled “A second name for a static site (the hosted service’s layout)”The hosted service runs the application on app.silencewatch.com and serves its
showcase site and documentation (site/) from the bare domain. The stack does
the same when it is given a second name, and does nothing about it when it is
not:
SILENCEWATCH_DOMAIN=app.example.com # the application — as aboveSILENCEWATCH_SITE_DOMAIN=example.com # optional: a static site on this nameBASE_URL=https://app.example.com # still the application's addressWith SILENCEWATCH_SITE_DOMAIN set, Caddy also serves the site volume on that
name, redirects www. to it, and sends the addresses the application used to
answer there (/p/…, /api/…, /login and the rest) to the application with a
308, so a bookmark, a verification email sent before the move or an old crontab
still lands. Unset, none of this exists: no second block, no redirect, and the
volume is never read. Both names must resolve to the swarm (add www. too), with
80 and 443 reachable.
The volume is filled by .github/workflows/site.yml on every change to site/
or docs/ that reaches main: each build is unpacked beside the others and a
symlink, live, is moved to it in one rename, so a visitor never sees half a
site. The last three are kept; going back is ln -sfn <older> live inside the
volume. That job connects to the same host as the release deploy, so it needs the
swarm to be a single node, or the SSH host to be the node Caddy is pinned to —
the volume is local to it, like Caddy’s certificates.
BIND_ADDRESS has no Swarm equivalent: the routing mesh publishes a port on
every node and cannot be told to bind one address. The application’s port is
still published there, because the deployment check reads /health through it,
so close it at the firewall — otherwise it is the same plain-HTTP bypass
BIND_ADDRESS closes under Compose:
sudo ufw allow 80,443/tcpsudo ufw deny 8080/tcpWho watches the watchman
Section titled “Who watches the watchman”SilenceWatch cannot monitor itself. A monitoring server that dies quietly is the one outage there is no recovering from: everything looks green because nothing is looking.
Set OUTBOUND_HEARTBEAT_URL to a dead man’s switch owned by someone else —
Healthchecks.io, Cronitor, or a colleague’s SilenceWatch:
OUTBOUND_HEARTBEAT_URL=https://hc-ping.com/<your-uuid>The server pings it every minute, but only while the detection loop has completed a pass recently. If detection stalls while HTTP keeps answering — the failure mode a naive uptime check misses entirely — the heartbeat stops and the external watchdog fires.
/health follows the same rule: it returns 503 when the detection loop has
stalled, so an orchestrator restarts a process that is technically alive and
practically useless.
Backups
Section titled “Backups”Everything lives in PostgreSQL. Nothing is stored on disk by the application.
docker compose exec db pg_dump -U silencewatch silencewatch | gzip > silencewatch-$(date +%F).sql.gzRestore:
gunzip -c silencewatch-2026-01-15.sql.gz | docker compose exec -T db psql -U silencewatch silencewatchAnd — this being what the product is for — monitor the backup job with a check.
Upgrading
Section titled “Upgrading”Edit the image tag in docker-compose.yml to the release you want, then:
docker compose pulldocker compose up -dThe tag is pinned rather than latest on purpose: a restart should not change
which version is running, and a bug report needs a version to name. Published
images are listed under
Releases.
Working from a checkout instead, with the build stanza uncommented:
git pulldocker compose build --pulldocker compose up -dMigrations are applied at startup. They are additive by design; when a release needs a destructive change, the notes say so and give the steps.
Docker Swarm, and upgrading without downtime
Section titled “Docker Swarm, and upgrading without downtime”Compose restarts the container, so an upgrade is a short outage. Swarm can start
the new version, wait for it to report healthy, and only then stop the old one.
docker-stack.yml in the repository root is that deployment.
It is not docker-compose.yml with a deploy: block added. docker stack deploy silently ignores depends_on, restart, and the short form of tmpfs,
so the two files differ where it matters — a read-only container whose /tmp
mount was dropped does not start at all.
What makes it seamless
Section titled “What makes it seamless”Two replicas, order: start-first, and the image’s own HEALTHCHECK. Swarm
brings a new task up, waits for /health, shifts traffic, and stops the old
one; with a single replica there is still a moment when the only healthy task is
the one being replaced.
Running two instances is safe here by design rather than by luck: the detection
loop and the alert queue claim rows with FOR UPDATE SKIP LOCKED, so they share
work instead of alerting twice, and the ingest cache is invalidated across
instances with LISTEN/NOTIFY.
The precondition is that migrations stay additive. During the changeover the old and new versions run against the same schema for a few seconds. That holds for every release whose notes do not say otherwise — when one does, deploy it as a brief planned outage instead.
If the new version never reports healthy, Swarm puts the old one back on its own
(failure_action: rollback), and the deploy job fails on the version check
rather than reporting a success that did not happen.
One-time setup
Section titled “One-time setup”docker swarm init # if it is not already onedocker node update --label-add silencewatch.db=true "$(docker node ls -q)"sudo install -d -m 750 /opt/silencewatchsudo cp .env /opt/silencewatch/.env # SECRET_KEY, POSTGRES_PASSWORD, BASE_URL…sudo chmod 600 /opt/silencewatch/.envThe node label is not optional. Without it a reschedule would start PostgreSQL
on another machine against an empty local volume — an instance that looks
healthy and has lost every check, ping and account. For the same reason the
database updates stop-first: two PostgreSQL processes on one data directory
corrupt it, so it takes a brief pause where the application does not.
Deploy by hand with:
set -a; . /opt/silencewatch/.env; set +aSILENCEWATCH_VERSION=0.1.0 docker stack deploy -c docker-stack.yml silencewatchDeploying every release automatically
Section titled “Deploying every release automatically”.github/workflows/release.yml runs the same deployment over SSH once a tag’s
image has been published and smoke tested. Tag a release, and the swarm is on it
a couple of minutes later without anyone logging in.
This is for a fork you operate. It is not needed to self-host: the manual
docker stack deploy above is the same operation, run when you choose.
1. A key that exists only for this
Section titled “1. A key that exists only for this”ssh-keygen -t ed25519 -C "github-actions@silencewatch" \ -f ~/.ssh/silencewatch_deploy -N ""-N "" means no passphrase, because CI cannot type one. That is exactly why
this is a dedicated key rather than yours: it can be revoked on its own, and it
never protected anything else.
2. A user for it on the manager node
Section titled “2. A user for it on the manager node”sudo adduser --disabled-password --gecos "" deploysudo usermod -aG docker deploysudo install -d -m 700 -o deploy -g deploy /home/deploy/.ssh
sudo tee /home/deploy/.ssh/authorized_keys <<EOFrestrict,pty $(cat ~/.ssh/silencewatch_deploy.pub)EOFsudo chown deploy:deploy /home/deploy/.ssh/authorized_keyssudo chmod 600 /home/deploy/.ssh/authorized_keysrestrict disables port forwarding, agent forwarding and X11, none of which a
deployment needs.
Membership of the
dockergroup is equivalent to root on that machine. The daemon runs as root and will mount any path on the host for you. This is inherent to deploying over SSH rather than a weakness of this setup, but it sets the value of the key: treat it as a root credential, keep it out of any other system, and revoke it by deleting the line fromauthorized_keys.
3. The four values
Section titled “3. The four values”| Secret | What to paste | Where it comes from |
|---|---|---|
SWARM_SSH_HOST |
the host alone, e.g. silencewatch.com |
no user@, no port |
SWARM_SSH_USER |
deploy |
the user created above |
SWARM_SSH_KEY |
the private key, in full | cat ~/.ssh/silencewatch_deploy |
SWARM_SSH_KNOWN_HOSTS |
the host’s fingerprints | ssh-keyscan -H <host> |
Two things that trip people up. SWARM_SSH_KEY is the file without the
.pub suffix — the .pub one belongs on the server — and it must include the
-----BEGIN/-----END lines. ssh-keyscan prints one line per key type: paste
all of them.
The last secret is what makes the connection safe. Without a pinned host key the
alternative is StrictHostKeyChecking=no, which hands the deploy key, and every
command that follows it, to whatever answers on that address. The job refuses to
run if the value is empty rather than falling back to trusting anything.
4. Where they go
Section titled “4. Where they go”Settings → Environments → production → Add environment secret.
Repository secrets work too, but scoping them to the environment means only the
job that targets it can read them — not the build job, and not a workflow added
later. production is also where a required reviewer goes if releases should
stop shipping unattended.
If the env file is not at /opt/silencewatch/.env, set a repository variable
(not a secret) named SWARM_ENV_FILE to its path. The values inside it stay on
the server: CI reads their names from the stack file and never sees them.
5. Check it before relying on it
Section titled “5. Check it before relying on it”ssh -i ~/.ssh/silencewatch_deploy deploy@<host> \ 'docker version --format "{{.Server.Version}}" && docker node ls'A version and a node list means the four secrets will work. A password prompt
means authorized_keys is not being read — usually its permissions, or those of
/home/deploy/.ssh.
The workflow connects on port 22. A different port needs -p adding to the two
ssh invocations in release.yml.
Diagnostics
Section titled “Diagnostics”For a support thread, produce a bundle:
docker compose exec silencewatch node packages/server/dist/diagnostics/support-bundle.js /tmp/bundle.jsondocker compose cp silencewatch:/tmp/bundle.json .It contains versions, the effective configuration as set/not-set rather than values, table sizes, index presence, migration state and health — no secrets, no ping keys, no email addresses. Read it before sending it anyway.
Scaling up
Section titled “Scaling up”Run more than one instance when a single process is no longer enough. Nothing
needs changing: the detection loop and the alert queue claim rows with
FOR UPDATE SKIP LOCKED, so instances share work without duplicating alerts, and
the ingestion cache is invalidated across instances through PostgreSQL
LISTEN/NOTIFY.
Beyond that the database is the limit. In order: raise INGEST_POOL_MAX, put the
ping table on faster storage, then lower PING_RETENTION_DAYS.