Job monitoring: get alerted when a job stops
Your jobs don’t tell you when they stop. Job monitoring answers a simple question: did my job run when it should, and did it finish properly? SilenceWatch answers it for any job that can send an HTTP request.
What is job monitoring?
Section titled “What is job monitoring?”A job is a task that runs without anyone starting it: a backup, an export, a billing run, an import, a purge, a batch, a periodic worker. It is also called a scheduled task, a batch or a cron job — cron is only one way of scheduling a job among others, which is why cron job monitoring is a special case. Job monitoring checks that each one runs at the expected time, in a normal amount of time and successfully, and alerts otherwise.
What you really watch in a job
Section titled “What you really watch in a job”Three things, and each fails differently:
- Presence. Did the job start at the expected time? It is the sneakiest failure, because it produces no error: a stopped scheduler, a deleted schedule, a container that no longer starts.
- Outcome. Did it finish successfully? A job that starts but crashes should be able to say so right away, with its exit code.
- Duration. Does it take the usual time? A process that doubles in duration often announces a problem before it stops.
Why other monitoring is not enough
Section titled “Why other monitoring is not enough”| What you have | What it sees | What it misses |
|---|---|---|
| Log alerts | The errors that were written | A job that does not run writes nothing |
| Uptime monitoring | That a service responds | That a scheduled job is running |
| APM / tracing | The behaviour of code that executes | The absence of execution |
| Heartbeat (SilenceWatch) | The arrival, or absence, of a signal at the expected time | The content of the work (you need an explicit failure signal) |
These tools are complementary. The heartbeat covers what none of the others sees: absence.
How SilenceWatch monitors a job
Section titled “How SilenceWatch monitors a job”Each job calls a ping URL at the end of its run and, if you like, at the start (to measure duration) and on failure:
GET|POST /p/<key> the run succeededGET|POST /p/<key>/start the run startedGET|POST /p/<key>/fail the run failedGET|POST /p/<key>/<code> exit code (0 succeeds)You declare the expected frequency (interval or cron expression, with time zone) and a grace period. Past the deadline, the check is late; once the grace period has passed, it is down, an incident opens and your channels are alerted (email, webhook, Slack, Microsoft Teams, Discord). The detail is in check states.
Which jobs to monitor
Section titled “Which jobs to monitor”- Database and file backups.
- Data imports and exports, ETL processing, synchronisations.
- Billing runs, reminders, bulk mailings.
- Purges, rotations, certificate renewals.
@Scheduledtasks and Quartz jobs of a Spring Boot application, discovered automatically.- The cron jobs of a server, a container or a Kubernetes CronJob.
For message queues or on-demand processing, a periodic heartbeat is not the right tool: SilenceWatch monitors jobs that must run at regular intervals.
Choosing a job monitoring tool
Section titled “Choosing a job monitoring tool”A few useful questions, whatever the tool:
- Does it detect absence? That is the heart of the matter: a tool that only looks at errors cannot see a job that no longer starts.
- Does it understand your schedules? Cron expressions, time zones, clock changes, a grace period.
- Does it measure duration, and tell a failure from lateness?
- Does it alert where you are, and can a channel be tested before the incident?
- Does it integrate without effort? One HTTP request for a script; one dependency for a Java application.
- Where does your data live? A hosted service, or a self-hostable tool that is never crippled.
Frequently asked questions
Section titled “Frequently asked questions”What is job monitoring?
It is watching that your jobs and scheduled tasks run: checking that each one runs at the expected time, in a normal amount of time and successfully, and alerting otherwise. Cron jobs are part of it.
How do I monitor a job that does not crash but no longer starts?
With a heartbeat: the job sends a signal on every run, and SilenceWatch alerts you if it does not arrive within the expected window. It is the only way to detect an absence, because a job that does not run produces no error.
What is the difference between job monitoring and log monitoring?
Log monitoring detects errors that were written. Heartbeat job monitoring detects that a job did not run at all, which a log cannot show since an absent job writes nothing.
Can I measure how long a job takes?
Yes. The job calls /start at the beginning and the result URL at the end; SilenceWatch measures the interval, or you can pass a duration yourself with the duration_ms parameter.
How do I get alerted when a job fails without waiting for the deadline?
Call /fail, or the URL with a non-zero exit code, when the job fails: the check goes down immediately, without waiting for the grace period to end.
Is SilenceWatch suitable for Java jobs?
Yes, with a Spring Boot starter that automatically discovers @Scheduled tasks and Quartz jobs and sends the heartbeats for you.