Monitor a cron backup, step by step
A backup nobody knows is running is not a backup. This guide shows how to monitor a cron backup so you hear about it if it does not start, fails or produces an empty file.
The three ways a backup fails
Section titled “The three ways a backup fails”- It stops starting (lost crontab, rebooted server, deleted user). Nothing is written anywhere: only a dead man’s switch detects it.
- It fails (database unreachable, disk full, changed password). The script can say so explicitly.
- It “succeeds” but produces an empty or truncated file. The exit code is fine, the backup is unusable. Only the script can check that before announcing success.
Step by step
Section titled “Step by step”-
Create the check. In SilenceWatch: name “PostgreSQL backup”, cron schedule
0 2 * * *with theEurope/Paristime zone, a grace period of one hour (more than the backup’s normal duration), environmentproduction. Copy the ping URL. -
Write a script that checks its own result.
/usr/local/bin/backup-db.sh #!/bin/bashset -uo pipefailURL=https://app.silencewatch.com/p/<ping-key>DUMP="/backups/db-$(date +%F).sql.gz"MIN_BYTES=1048576 # below this, the backup is suspectcurl -fsS -m 10 --retry 3 "$URL/start" || trueif pg_dump -U app mydb | gzip > "$DUMP" \&& [ "$(stat -c %s "$DUMP")" -ge "$MIN_BYTES" ]; thencurl -fsS -m 10 --retry 3 --data-raw "OK $(basename "$DUMP") $(du -h "$DUMP" | cut -f1)" "$URL" || trueelsecurl -fsS -m 10 --retry 3 --data-raw "Backup failed or too small: $DUMP" "$URL/fail" || trueexit 1fiThe start call (
/start) lets SilenceWatch measure the duration; the ping body keeps the file size in the history; failure calls/failand the check goes down immediately, without waiting for the deadline. The|| truekeep a failed ping from failing the backup itself. -
Schedule it with cron.
crontab -e 0 2 * * * /usr/local/bin/backup-db.sh -
Test the three cases. Run the script by hand: the check goes UP. Break database access and run it again: it goes down and the alert is sent. Disable the crontab line: at the next deadline plus the grace period it goes down too. That is the three failures covered.
-
Add an alert channel if you have not, and test it: see alert channels.
Going further
Section titled “Going further”- Test your restores. Monitoring the backup proves it is made, not that it restores. Plan a periodic restore, monitored by its own check too.
- One check per backup. Database, files, objects: one check each, so you know which is missing.
- Watch the duration. A backup that doubles in duration often announces a filling disk or a database growing too fast.
For other environments (systemd, Docker, Kubernetes, GitHub Actions), see the cron monitoring examples.
Frequently asked questions
Section titled “Frequently asked questions”How do I get alerted if a cron backup did not run?
Have it send a ping to SilenceWatch at the end of every successful backup. If the ping does not arrive by the expected time plus the grace period, the check goes down and your alert channels are notified.
How do I detect an empty backup?
Have the script check the file size before announcing success, and call the /fail URL when it is below a reasonable threshold.
What grace period should I choose for a backup?
A little more than its normal duration, with margin: one to two hours for a twenty-minute backup avoids false alerts without delaying the news needlessly.