Skip to content

Monitor a cron backup, step by step

A backup nobody knows is running is not a backup. This guide shows how to monitor a cron backup so you hear about it if it does not start, fails or produces an empty file.

  1. It stops starting (lost crontab, rebooted server, deleted user). Nothing is written anywhere: only a dead man’s switch detects it.
  2. It fails (database unreachable, disk full, changed password). The script can say so explicitly.
  3. It “succeeds” but produces an empty or truncated file. The exit code is fine, the backup is unusable. Only the script can check that before announcing success.
  1. Create the check. In SilenceWatch: name “PostgreSQL backup”, cron schedule 0 2 * * * with the Europe/Paris time zone, a grace period of one hour (more than the backup’s normal duration), environment production. Copy the ping URL.

  2. Write a script that checks its own result.

    /usr/local/bin/backup-db.sh
    #!/bin/bash
    set -uo pipefail
    URL=https://app.silencewatch.com/p/<ping-key>
    DUMP="/backups/db-$(date +%F).sql.gz"
    MIN_BYTES=1048576 # below this, the backup is suspect
    curl -fsS -m 10 --retry 3 "$URL/start" || true
    if pg_dump -U app mydb | gzip > "$DUMP" \
    && [ "$(stat -c %s "$DUMP")" -ge "$MIN_BYTES" ]; then
    curl -fsS -m 10 --retry 3 --data-raw "OK $(basename "$DUMP") $(du -h "$DUMP" | cut -f1)" "$URL" || true
    else
    curl -fsS -m 10 --retry 3 --data-raw "Backup failed or too small: $DUMP" "$URL/fail" || true
    exit 1
    fi

    The start call (/start) lets SilenceWatch measure the duration; the ping body keeps the file size in the history; failure calls /fail and the check goes down immediately, without waiting for the deadline. The || true keep a failed ping from failing the backup itself.

  3. Schedule it with cron.

    crontab -e
    0 2 * * * /usr/local/bin/backup-db.sh
  4. Test the three cases. Run the script by hand: the check goes UP. Break database access and run it again: it goes down and the alert is sent. Disable the crontab line: at the next deadline plus the grace period it goes down too. That is the three failures covered.

  5. Add an alert channel if you have not, and test it: see alert channels.

  • Test your restores. Monitoring the backup proves it is made, not that it restores. Plan a periodic restore, monitored by its own check too.
  • One check per backup. Database, files, objects: one check each, so you know which is missing.
  • Watch the duration. A backup that doubles in duration often announces a filling disk or a database growing too fast.

For other environments (systemd, Docker, Kubernetes, GitHub Actions), see the cron monitoring examples.

How do I get alerted if a cron backup did not run?

Have it send a ping to SilenceWatch at the end of every successful backup. If the ping does not arrive by the expected time plus the grace period, the check goes down and your alert channels are notified.

How do I detect an empty backup?

Have the script check the file size before announcing success, and call the /fail URL when it is below a reasonable threshold.

What grace period should I choose for a backup?

A little more than its normal duration, with margin: one to two hours for a twenty-minute backup avoids false alerts without delaying the news needlessly.