Schedules, grace period and check states
A check has two settings that decide everything: when the next heartbeat is expected (the schedule) and how late you tolerate (the grace period).
The schedule
Section titled “The schedule”Two forms, never both:
- An interval: “every N seconds”. The minimum is 30 seconds. Suited to regular processing (every 5 minutes, every hour).
- A cron expression: with 5 fields (Unix) or 6 (leading seconds, as Spring and Quartz write them), and an IANA time zone (
Europe/Paris). The forms?,L,5LandMON#2are accepted. Quartz’sW(nearest weekday) is rejected: the server could not compute the deadline, and a check whose deadline cannot be computed is worse than no check.
The time zone matters for cron expressions: 0 2 * * * is not the same moment in Paris and in New York, nor before and after a clock change.
The grace period
Section titled “The grace period”It is the lateness accepted after the deadline before you are alerted. Size it to the job: a backup that normally takes twenty minutes deserves a grace period of more than twenty minutes. Too short and you get false alerts; too long and you hear late.
The states
Section titled “The states”| State | What it means |
|---|---|
| Waiting | No ping has arrived yet. |
| UP | The last ping arrived on time. |
| Late | The deadline has passed, but the grace period is still running. No alert is sent. |
| Down | The deadline and the grace period have passed, or the job reported a failure (/fail, or a non-zero exit code). An incident opens and every enabled channel is alerted. |
| Paused | You suspended the check: it does not alert and its pings are ignored. |
Colour is never the only signal: each state also has a label and a shape, so it stays readable in greyscale.
Incidents and recovery
Section titled “Incidents and recovery”When a check goes down, an incident opens and the channels are alerted. The next ping brings it UP, closes the incident and sends a recovery alert, only to those who were told about the outage. The history keeps each incident’s duration and the number of alerts sent.
Detection runs every 10 seconds; each pass costs only what the late checks cost, whatever the total number of checks.
Automatically declared checks
Section titled “Automatically declared checks”Checks created by a starter are marked auto. If the application stops declaring a job (renamed, removed), its check becomes orphaned: it is flagged, never deleted automatically, because an automatic delete would destroy the history at the first refactoring. Deleting it is your decision.