Dead man's switch: definition and use
A dead man’s switch triggers an action when a regular signal stops arriving. Applied to jobs, it is the reliable way to detect that a scheduled task no longer runs.
The principle
Section titled “The principle”The name comes from the safety controls of trains and machinery: the operator must hold or press a device regularly, and it is the stopping of the gesture, not a gesture, that raises the alarm. In IT it is also called heartbeat or check-in: a program sends a signal at regular intervals, and a service watches that the signal arrives.
If the signal arrives on time, all is well and nothing happens. If it is missing, an alert goes out.
Why it is the right approach for scheduled jobs
Section titled “Why it is the right approach for scheduled jobs”The usual kinds of monitoring look for an event: an error in a log, a failing HTTP response, an abnormal trace. But the most common failure of a scheduled task is that it stops starting, and an absence produces no event to detect. The dead man’s switch reverses the logic: you do not look for the problem, you require proof that everything went well.
That is also what makes it simple to set up: one HTTP request at the end of the job, and nothing to install on the machine.
What a dead man’s switch detects
Section titled “What a dead man’s switch detects”- a stopped cron or scheduler, a deleted or overwritten schedule;
- a machine or container that is off, a deployment that removed the job;
- a job that hangs or never reaches its end;
- a job that fails and says so explicitly (
/fail), or that becomes abnormally long.
Its limits
Section titled “Its limits”A dead man’s switch attests to execution, not to the quality of the work: a job can run, send its signal and still produce a wrong result. For a backup, the script must therefore check its own result before sending the success signal, as in the backup monitoring guide. It also suits periodic tasks: for a message queue or on-demand processing, other monitoring is needed. Finally, it is “passive” monitoring: the job calls in, the service does not poll.
SilenceWatch, a dead man’s switch for your jobs
Section titled “SilenceWatch, a dead man’s switch for your jobs”SilenceWatch is a complete implementation: one ping URL per check, a schedule (interval or cron with time zone), a grace period, incidents, alerts by email, webhook, Slack, Microsoft Teams or Discord, and a history of every signal. It is open source (Apache 2.0) and can be self-hosted.
Frequently asked questions
Section titled “Frequently asked questions”What is a dead man’s switch in IT?
It is a mechanism that raises an alert when a regular signal stops arriving. A program sends a heartbeat at regular intervals, and a service watches for it. The absence of the signal is what triggers the alert.
What is the difference between a dead man’s switch and heartbeat monitoring?
They are two names for the same idea. Heartbeat monitoring refers to the regular signal the program sends; dead man’s switch stresses that it is its stopping that raises the alarm.
Does a dead man’s switch detect a job that runs but produces a wrong result?
Not by itself: it attests to execution. Have the script check its result before sending the success signal, and call the failure URL when something is wrong.