Monitor a Kubernetes CronJob
A Kubernetes CronJob can stop producing Jobs without a single Pod failing: there is then nothing to observe. A ping sent at the end of the container turns that absence into an alert.
What Kubernetes does not tell you
Section titled “What Kubernetes does not tell you”A CronJob creates a Job at each deadline, and the Job creates a Pod. Visible failures (a Pod in error, Failed) are watched with the usual tools. But several situations produce none:
- the CronJob is suspended (
spec.suspend: true) and stays so because nobody remembers; - the concurrency policy (
concurrencyPolicy) skips a run: withForbid, if the previous one is still running, the new one is not created; - a deadline is missed beyond
startingDeadlineSeconds(cluster unavailable, controller restarted); - the time zone is not what you thought: the
spec.timeZonefield exists since Kubernetes 1.27, and without it the controller’s time applies; - the whole cluster or its controller is down.
In all these cases no Pod is created, so no Pod fails. That is exactly what a dead man’s switch detects.
Adding the ping to the CronJob
Section titled “Adding the ping to the CronJob”The ping URL is a secret: store it in a Secret, and call it at the end of the container, on success only.
apiVersion: batch/v1kind: CronJobmetadata: name: backupspec: schedule: "0 2 * * *" timeZone: "Europe/Paris" concurrencyPolicy: Forbid jobTemplate: spec: template: spec: restartPolicy: OnFailure containers: - name: backup image: my-image:latest command: ["/bin/sh", "-c"] args: - | curl -fsS -m 10 --retry 3 "$PING_URL/start" || true /app/backup.sh code=$? curl -fsS -m 10 --retry 3 "$PING_URL/$code" || true exit $code env: - name: PING_URL valueFrom: secretKeyRef: name: silencewatch key: ping-urlThe container calls /start at the beginning, then the URL suffixed with the exit code at the end: 0 succeeds, any other code takes the check down immediately. The image must contain curl; otherwise add it or use a second container.
Check settings
Section titled “Check settings”- Schedule: the same cron expression as
spec.schedule, with the same time zone. - Grace period: more than the job’s normal duration, plus the usual Pod start-up latency.
- One check per CronJob, with the environment (
production,staging) so clusters do not get mixed up.
See also the examples for other environments and check states.
Frequently asked questions
Section titled “Frequently asked questions”How do I know a Kubernetes CronJob did not run?
Have the container send a ping at the end and watch for it with SilenceWatch: if the ping is missing at the deadline plus the grace period you are alerted, whether the Pod failed or was never created.
Why does a CronJob no longer create Jobs?
Common causes: the CronJob is suspended, concurrencyPolicy Forbid skips the run while the previous one is still running, a deadline was missed beyond startingDeadlineSeconds, or the time zone is not the one expected.
Does the ping fail my Job if it is unreachable?
Not if you add || true after the curl: a lost ping then does not change the job’s exit code.