Skip to content

Monitor a Kubernetes CronJob

A Kubernetes CronJob can stop producing Jobs without a single Pod failing: there is then nothing to observe. A ping sent at the end of the container turns that absence into an alert.

A CronJob creates a Job at each deadline, and the Job creates a Pod. Visible failures (a Pod in error, Failed) are watched with the usual tools. But several situations produce none:

  • the CronJob is suspended (spec.suspend: true) and stays so because nobody remembers;
  • the concurrency policy (concurrencyPolicy) skips a run: with Forbid, if the previous one is still running, the new one is not created;
  • a deadline is missed beyond startingDeadlineSeconds (cluster unavailable, controller restarted);
  • the time zone is not what you thought: the spec.timeZone field exists since Kubernetes 1.27, and without it the controller’s time applies;
  • the whole cluster or its controller is down.

In all these cases no Pod is created, so no Pod fails. That is exactly what a dead man’s switch detects.

The ping URL is a secret: store it in a Secret, and call it at the end of the container, on success only.

cronjob.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
name: backup
spec:
schedule: "0 2 * * *"
timeZone: "Europe/Paris"
concurrencyPolicy: Forbid
jobTemplate:
spec:
template:
spec:
restartPolicy: OnFailure
containers:
- name: backup
image: my-image:latest
command: ["/bin/sh", "-c"]
args:
- |
curl -fsS -m 10 --retry 3 "$PING_URL/start" || true
/app/backup.sh
code=$?
curl -fsS -m 10 --retry 3 "$PING_URL/$code" || true
exit $code
env:
- name: PING_URL
valueFrom:
secretKeyRef:
name: silencewatch
key: ping-url

The container calls /start at the beginning, then the URL suffixed with the exit code at the end: 0 succeeds, any other code takes the check down immediately. The image must contain curl; otherwise add it or use a second container.

  • Schedule: the same cron expression as spec.schedule, with the same time zone.
  • Grace period: more than the job’s normal duration, plus the usual Pod start-up latency.
  • One check per CronJob, with the environment (production, staging) so clusters do not get mixed up.

See also the examples for other environments and check states.

How do I know a Kubernetes CronJob did not run?

Have the container send a ping at the end and watch for it with SilenceWatch: if the ping is missing at the deadline plus the grace period you are alerted, whether the Pod failed or was never created.

Why does a CronJob no longer create Jobs?

Common causes: the CronJob is suspended, concurrencyPolicy Forbid skips the run while the previous one is still running, a deadline was missed beyond startingDeadlineSeconds, or the time zone is not the one expected.

Does the ping fail my Job if it is unreachable?

Not if you add || true after the curl: a lost ping then does not change the job’s exit code.