Cron jobs are the quiet workhorses of web automation, but when a scheduled task runs twice at the same moment, the results can be messy: duplicate emails, double charges, or corrupted data. Preventing duplicate execution is essential for any website owner who relies on scheduled tasks. This guide explains practical techniques to keep your cron jobs honest, whether you run a small blog or a busy application. For more automation insights, the FreeCronJob blog offers regular updates on reliable scheduling.
File Locks: The First Line of Defense
A cron lock using a file is the simplest approach. The script creates a lock file before starting and removes it when finished. If the lock file already exists, the script exits immediately. This works well on a single server. Use flock on Linux systems, which releases the lock automatically if the process dies, avoiding a stale lock that blocks future runs.
Database Locks for Multi-Server Setups
When your application runs on several servers, a file lock on one machine cannot protect the others. A database lock solves this. Create a table with a unique key representing the job name. Before running, insert a row with that key. If the insert fails because the key already exists, another instance is running. Release the lock by deleting the row after completion. Add a timestamp and a timeout so a crashed job does not leave a permanent lock.
Idempotency Keys: Make Duplicates Harmless
Sometimes a duplicate execution is unavoidable. Idempotency makes the second run a no-op. For example, if a job sends a welcome email, store a unique idempotency key in the database for each recipient. Before sending, check whether the key exists. If it does, skip the action. This technique is powerful because it protects against duplicates even when locks fail. Many payment systems use idempotency keys to prevent double charges.
Safe Recovery After Interrupted Jobs
A job can crash mid-run, leaving locks behind. Always design your locks with expiration. For file locks, use flock with a timeout. For database locks, store a started_at timestamp and allow a new run to take over after a threshold. Also, write your job steps to be resumable. Process items in batches and record progress in a table, so a restarted job picks up where it left off instead of repeating everything.
| Method | Best For | Stale Lock Risk | Recovery |
|---|---|---|---|
| File lock | Single server | Low with flock | Automatic on process exit |
| Database lock | Multi-server | Medium without timeout | Delete row or timeout |
| Idempotency key | Duplicate-sensitive actions | None | Safe re-run by design |
Choosing the right protection depends on your infrastructure. A small site may only need a file cron lock, while a distributed app benefits from database locks and idempotency keys. Combine these techniques to build a robust schedule and keep your automation reliable.
