Skip to content

Job monitoring: get alerted when a job stops

Your jobs don’t tell you when they stop. Job monitoring answers a simple question: did my job run when it should, and did it finish properly? SilenceWatch answers it for any job that can send an HTTP request.

A job is a task that runs without anyone starting it: a backup, an export, a billing run, an import, a purge, a batch, a periodic worker. It is also called a scheduled task, a batch or a cron job — cron is only one way of scheduling a job among others, which is why cron job monitoring is a special case. Job monitoring checks that each one runs at the expected time, in a normal amount of time and successfully, and alerts otherwise.

Three things, and each fails differently:

  • Presence. Did the job start at the expected time? It is the sneakiest failure, because it produces no error: a stopped scheduler, a deleted schedule, a container that no longer starts.
  • Outcome. Did it finish successfully? A job that starts but crashes should be able to say so right away, with its exit code.
  • Duration. Does it take the usual time? A process that doubles in duration often announces a problem before it stops.
What you have What it sees What it misses
Log alerts The errors that were written A job that does not run writes nothing
Uptime monitoring That a service responds That a scheduled job is running
APM / tracing The behaviour of code that executes The absence of execution
Heartbeat (SilenceWatch) The arrival, or absence, of a signal at the expected time The content of the work (you need an explicit failure signal)

These tools are complementary. The heartbeat covers what none of the others sees: absence.

Each job calls a ping URL at the end of its run and, if you like, at the start (to measure duration) and on failure:

GET|POST /p/<key> the run succeeded
GET|POST /p/<key>/start the run started
GET|POST /p/<key>/fail the run failed
GET|POST /p/<key>/<code> exit code (0 succeeds)

You declare the expected frequency (interval or cron expression, with time zone) and a grace period. Past the deadline, the check is late; once the grace period has passed, it is down, an incident opens and your channels are alerted (email, webhook, Slack, Microsoft Teams, Discord). The detail is in check states.

  • Database and file backups.
  • Data imports and exports, ETL processing, synchronisations.
  • Billing runs, reminders, bulk mailings.
  • Purges, rotations, certificate renewals.
  • @Scheduled tasks and Quartz jobs of a Spring Boot application, discovered automatically.
  • The cron jobs of a server, a container or a Kubernetes CronJob.

For message queues or on-demand processing, a periodic heartbeat is not the right tool: SilenceWatch monitors jobs that must run at regular intervals.

A few useful questions, whatever the tool:

  1. Does it detect absence? That is the heart of the matter: a tool that only looks at errors cannot see a job that no longer starts.
  2. Does it understand your schedules? Cron expressions, time zones, clock changes, a grace period.
  3. Does it measure duration, and tell a failure from lateness?
  4. Does it alert where you are, and can a channel be tested before the incident?
  5. Does it integrate without effort? One HTTP request for a script; one dependency for a Java application.
  6. Where does your data live? A hosted service, or a self-hostable tool that is never crippled.
What is job monitoring?

It is watching that your jobs and scheduled tasks run: checking that each one runs at the expected time, in a normal amount of time and successfully, and alerting otherwise. Cron jobs are part of it.

How do I monitor a job that does not crash but no longer starts?

With a heartbeat: the job sends a signal on every run, and SilenceWatch alerts you if it does not arrive within the expected window. It is the only way to detect an absence, because a job that does not run produces no error.

What is the difference between job monitoring and log monitoring?

Log monitoring detects errors that were written. Heartbeat job monitoring detects that a job did not run at all, which a log cannot show since an absent job writes nothing.

Can I measure how long a job takes?

Yes. The job calls /start at the beginning and the result URL at the end; SilenceWatch measures the interval, or you can pass a duration yourself with the duration_ms parameter.

How do I get alerted when a job fails without waiting for the deadline?

Call /fail, or the URL with a non-zero exit code, when the job fails: the check goes down immediately, without waiting for the grace period to end.

Is SilenceWatch suitable for Java jobs?

Yes, with a Spring Boot starter that automatically discovers @Scheduled tasks and Quartz jobs and sends the heartbeats for you.