Skip to content

Schedules, grace period and check states

A check has two settings that decide everything: when the next heartbeat is expected (the schedule) and how late you tolerate (the grace period).

Two forms, never both:

  • An interval: “every N seconds”. The minimum is 30 seconds. Suited to regular processing (every 5 minutes, every hour).
  • A cron expression: with 5 fields (Unix) or 6 (leading seconds, as Spring and Quartz write them), and an IANA time zone (Europe/Paris). The forms ?, L, 5L and MON#2 are accepted. Quartz’s W (nearest weekday) is rejected: the server could not compute the deadline, and a check whose deadline cannot be computed is worse than no check.

The time zone matters for cron expressions: 0 2 * * * is not the same moment in Paris and in New York, nor before and after a clock change.

It is the lateness accepted after the deadline before you are alerted. Size it to the job: a backup that normally takes twenty minutes deserves a grace period of more than twenty minutes. Too short and you get false alerts; too long and you hear late.

State What it means
Waiting No ping has arrived yet.
UP The last ping arrived on time.
Late The deadline has passed, but the grace period is still running. No alert is sent.
Down The deadline and the grace period have passed, or the job reported a failure (/fail, or a non-zero exit code). An incident opens and every enabled channel is alerted.
Paused You suspended the check: it does not alert and its pings are ignored.

Colour is never the only signal: each state also has a label and a shape, so it stays readable in greyscale.

When a check goes down, an incident opens and the channels are alerted. The next ping brings it UP, closes the incident and sends a recovery alert, only to those who were told about the outage. The history keeps each incident’s duration and the number of alerts sent.

Detection runs every 10 seconds; each pass costs only what the late checks cost, whatever the total number of checks.

Checks created by a starter are marked auto. If the application stops declaring a job (renamed, removed), its check becomes orphaned: it is flagged, never deleted automatically, because an automatic delete would destroy the history at the first refactoring. Deleting it is your decision.