Job Scheduling and Notifications
The job system provides asynchronous task execution with scheduling, worker management, and notification delivery. Jobs are the primary mechanism for background processing — anything from score recalculation to data imports to webhook delivery.
Architecture
Programs
A program defines a type of work that can be executed. Programs are registered during tenant setup and referenced by jobs.
| Attribute | Purpose |
|---|---|
| Code | Unique identifier within the tenant |
| Name | Display name with i18n support |
| Priority | Default priority — higher runs first |
| Parameters | Schema defining accepted parameters |
| Timeout | Maximum execution time |
Each program has one or more actions that are invoked when the program executes. This separates the "what to do" (program) from the "how to run it" (actions).
Worker management
Worker managers represent processing pools, each with a configurable concurrency limit.
Workers are individual processing instances registered under a manager. Workers claim jobs and report their status.
Jobs
A job is a runtime instance of a program. When a user or system submits work, a job record is created.
Lifecycle
| Phase | Status | Meaning |
|---|---|---|
PENDING | NORMAL | Waiting to be picked up by a worker; scheduled jobs start here |
STARTING | NORMAL | Being picked up by a worker |
RUNNING | NORMAL | Actively executing |
COMPLETING | NORMAL | Finishing up |
COMPLETED | NORMAL | Finished successfully |
COMPLETED | ERROR | Finished with errors |
COMPLETED | WARNING | Finished with warnings |
The platform dispatcher creates scheduled jobs in the PENDING phase. From there they progress through the same lifecycle as any other job.
Key attributes
| Attribute | Purpose |
|---|---|
| Job number | Auto-incrementing identity |
| Program | The program this job runs |
| Parameters | Input parameters for this execution |
| Scheduled start | When the job should start (default: now) |
| Schedule | The job schedule that created this job, if any |
| Parent job | For child jobs (job chaining) |
| Hold flag | Prevents the job from being picked up |
| Terminated flag | Marks the job as cancelled |
User context
Jobs capture the submitting user's context at creation time. This allows the job to execute with the same permissions as the user who submitted it.
For scheduled jobs, the job runs attributed to the schedule owner using the platform service identity. No user tokens are stored on scheduler-created job rows — delivery credentials are minted at outbox time from the schedule's owning user via the platform service identity mechanism.
Job schedules
A job schedule is a self-contained recurring work definition. It specifies the program to run, the parameters to pass, and its own timetable. The platform dispatcher reads due schedules every minute and creates jobs directly — no run depends on the previous run completing successfully.
Each schedule holds:
| Attribute | Purpose |
|---|---|
| Program code | The program to execute (same as a job's program field) |
| Parameters | Input parameters passed to each created job |
| Cron expression | Minute, hour, day-of-month, month, day-of-week fields in standard cron syntax |
| Jitter factor | Optional random offset applied to the minute field to smooth load distribution |
next_fire_date | Precomputed timestamp of the next due fire — validated and set by the platform at write time |
active | Whether the schedule is enabled (pause/resume; see lifecycle below) |
end_date | Optional date after which the schedule stops firing |
| Notification group | Health-event routing; overrides the program's notification group when set |
Schedule lifecycle
Creating a schedule — the cron expression is validated at save time. If the expression is malformed or cannot produce a next fire date before end_date, the write is rejected. The initial next_fire_date is computed and stored immediately.
Pausing and resuming — set active = false to pause. The schedule stops producing jobs until reactivated. On resumption, next_fire_date is recomputed from the current time — missed fires are not caught up. This is intentional: catch-up runs would process stale data and produce confusing results.
Retiring a schedule — set end_date to a past or future date. Once the date passes, the dispatcher stops firing the schedule. The schedule record and all historical run data are retained.
Overlap policy
The dispatcher checks for an active (non-terminal) job for each schedule before firing. If a previous run is still in PENDING, STARTING, RUNNING, or COMPLETING, the dispatcher skips that fire rather than creating a second overlapping run. Skips are tracked (consecutive_skips) and the schedule health is updated accordingly.
Schedule health and notifications
Health states
The dispatcher maintains a health_status on each schedule:
| State | Meaning |
|---|---|
OK | Schedule is firing normally |
SKIPPING | Consecutive skips have crossed the threshold — the previous run keeps holding the slot |
FAILING | Recent runs have completed with ERROR status |
EXPIRED | The schedule's end_date has passed or no future fire is possible before end_date |
Schedule Health view
The Schedule Health view in the Platform menu (Job user) shows the current health_status alongside last run phase/status, run counts, and error counts for the past 24 hours. Use this view to identify schedules that need attention without querying job history directly.
Notification events
Users subscribe to schedule health events by configuring a notification group on the schedule. The schedule-level group overrides the program's notification group, allowing different schedules running the same program to route alerts to different teams.
Health events delivered via notification groups:
| Event | Trigger |
|---|---|
| Run failed | A job created by this schedule completes with ERROR status |
| Schedule repeatedly skipping | Consecutive skips cross the configured threshold (default: 3) |
| Schedule expired | The schedule's end_date passes — one final notification is sent |
Notifications use the existing Liquid template path: the send_job_notifications trigger routes events through the notification outbox to email delivery. No separate notification infrastructure is required.
Job logs
Structured log entries for job execution, including severity level and human-readable message.