← blog · September 20, 2026

Why Scheduled Jobs Break Quietly: Missed Runs, Overlap and Single Instance in cron, systemd Timers and Kubernetes CronJobs

Scheduled jobs break quietly in three ways: the missed run, the overlapping run and the failure nobody sees. How cron, systemd timers and Kubernetes CronJobs answer each question, which fields to set, and which traps to expect.

Writing a scheduled job takes five minutes. Making sure that job runs at the right time, exactly once, and visibly, every day for a year is a separate engineering problem. Backups, report generation, certificate renewal and queue cleanup are the usual suspects, and most of them are eventually discovered with the sentence "it turns out this has not run for months". This article compares three schedulers (classic cron, systemd timers and Kubernetes CronJobs) against the same three questions: what happens to a missed run, what happens to an overlapping run, and who notices a failed run.

Three ways to fail quietly

Scheduled jobs break silently in three distinct ways.

The missed run. The machine was off, the scheduler process was restarting, or the cluster controller was busy. The window passed and the job never started. Nobody knows until the next window, and for a daily backup at 04:00 that means one day of data is simply absent.

The overlapping run. The job normally takes two minutes, but one day the data grew and it took thirty. A scheduler that fires every five minutes started six more instances in the meantime. Six processes writing to the same file, locking the same table, or sending the same request to the same API. The outcome is corrupted data or a system collapsing under load it created for itself.

The silent failure. The job started, died on the first line with exit code 126 or 127, and its output went somewhere nobody reads. The scheduler did its work; the job did none.

A good scheduling setup has an explicit answer to each of the three. A setup without answers is not a working setup, it is a setup that has not broken yet.

Classic cron: the most common and the weakest

cron promises nothing beyond triggering. It does not catch up on missed runs; if the machine shut down at 03:55 and came back at 04:05, the 04:00 job is ignored. anacron closes that gap only at daily, weekly and monthly granularity, which is useless for anything measured in minutes. There is no overlap protection at all; every trigger starts a fresh process. Failure reporting depends on local mail, and on most servers local mail is not configured, so the output vanishes.

Then there is the environment problem. cron jobs run with an almost empty PATH and usually under /bin/sh. A command that works in your interactive shell dies under cron with "command not found" and nobody sees it. Exit codes 126 (file exists but is not executable) and 127 (command not found) are the classic silent killers of scheduled work; a deployment that drops the execute bit on a script is enough.

If you must keep cron, do at least these three things: put set -euo pipefail and an absolute PATH at the top of the script, wrap the command in flock, and redirect output to a file or log shipper with something that alerts on failure.

*/5 * * * * flock -n /run/lock/report.lock /opt/jobs/report.sh >> /var/log/report.log 2>&1

flock -n exits immediately if the lock cannot be taken, so a new run never starts while the previous one is still going. That single line removes most of the overlap problem.

systemd timers: the right default on a single host

systemd timers answer all three questions natively, which is why they should replace cron on any single server.

For missed runs there is Persistent=true: a trigger that passed while the machine was off is made up once at boot. For overlap you do nothing special; the timer activates a service unit, and if that unit is still active a second instance is not started. Failures show up in systemctl --failed and the journal, and OnFailure= can hand off to another unit, for example one that sends a notification.

# /etc/systemd/system/report.service
[Unit]
Description=Daily report
OnFailure=notify@%n.service

[Service]
Type=oneshot
ExecStart=/opt/jobs/report.sh
TimeoutStartSec=30min
# /etc/systemd/system/report.timer
[Timer]
OnCalendar=*-*-* 04:00:00
Persistent=true
RandomizedDelaySec=10min

[Install]
WantedBy=timers.target

Type=oneshot keeps the service "active" until it finishes; overlap protection depends on that. TimeoutStartSec kills a stuck job, and without it a single hung run swallows every subsequent trigger forever. RandomizedDelaySec stops dozens of servers set to the same minute from hitting a shared backend in the same second. systemctl list-timers shows, in one table, when each timer last fired and when it fires next; cron has no equivalent.

Kubernetes CronJob: the only option in a cluster, with its own traps

In Kubernetes the same three questions map to CronJob fields.

apiVersion: batch/v1
kind: CronJob
metadata:
  name: report
spec:
  schedule: "0 4 * * *"
  timeZone: "Europe/Istanbul"
  concurrencyPolicy: Forbid
  startingDeadlineSeconds: 600
  successfulJobsHistoryLimit: 3
  failedJobsHistoryLimit: 5
  jobTemplate:
    spec:
      backoffLimit: 2
      activeDeadlineSeconds: 1800
      template:
        spec:
          restartPolicy: Never
          containers:
            - name: report
              image: registry.example.com/report:1.4.2

concurrencyPolicy is the overlap answer. Forbid skips the new trigger if the previous Job is still running, Replace kills the old one and starts the new one, and the default Allow runs them side by side exactly like cron. For any job that is not idempotent, Forbid is the right choice.

startingDeadlineSeconds is the missed-run answer and it cuts both ways. If the controller can start a trigger within this window it starts it, even if late; past the window the run counts as missed. If the field is unset and the controller finds it has missed more than one hundred schedules since the last scheduled time, it refuses to start the Job at all and only records an event. That is exactly how a CronJob that was left on suspend: true for a long time, or whose controller was down for a while, ends up "enabled but never running".

activeDeadlineSeconds cuts off a stuck run and backoffLimit sets how many retries happen. With restartPolicy: Never each retry is a separate Pod and the logs of the failed attempt remain inspectable; OnFailure restarts the same Pod and erases the trace of the previous attempt. Once the history limits are exceeded, Jobs are deleted; keeping failed Jobs around a little longer is the only way to answer "why did it not run last night".

Single instance across several machines

If a systemd timer or cron entry runs the same job on three servers, you get three instances. flock only locks the local filesystem. For a single instance across machines the lock has to live somewhere shared. The cheapest and most robust option is an advisory lock in the database the job already talks to:

SELECT pg_try_advisory_lock(7231);

In PostgreSQL this returns true if the lock was acquired and false if another session holds it, and the lock is released automatically when the session ends. The script exits quietly on false. This has far fewer moving parts than standing up a separate coordination service just for leader election.

Traps

Time zones. cron and CronJob default to the time zone of the host or the controller. Containers are UTC, hosts are local time, and the team thinks in local time. A job set to 02:30 runs twice in some years and never in others around daylight saving transitions. Set timeZone explicitly on CronJobs, add the zone to OnCalendar in systemd, or put everything on UTC.

Catch-up is dangerous on its own. Persistent=true or a long startingDeadlineSeconds means "run the missed job now". That may be the wrong behaviour for the missed window: a 04:00 backup that runs at 11:00 lands in the middle of production load. Unless the job contains logic like "if the missed window is this far in the past, skip" or "run, but behave differently", leave catch-up off and turn the miss into an alert instead.

Silence is the default for failure. None of the three schedulers tells anyone on its own. An alert of the form "no successful run for this long" is mandatory and must live independently of the job. If the job reports its own success, a job that never starts never reports; so the alert must fire on "the success report did not arrive", not on "a failure report arrived".

The scheduler and the job are not the same thing. A healthy scheduler does not mean the job ran. The exit code 126 example above is exactly this: every trigger succeeded, every job failed. Monitoring has to look at the outcome of the job.

When not to use a scheduler at all

If you need sub-minute triggering, if the run count is high and each unit of work is short, or if the work has to be spread across several workers, this is not a scheduling problem but a queue problem. Use the scheduler only to produce work for the queue. Likewise, do not satisfy "run when this event happens" with "check every minute"; short-interval polling adds latency and load and multiplies every trap listed above.

The selection rule is short: single host, use a systemd timer; cluster, use a CronJob with concurrencyPolicy: Forbid and startingDeadlineSeconds set; several hosts needing one instance, use either of those plus a database lock. Classic cron belongs only on legacy systems you cannot touch, and even there it should not be left without flock and output redirection.