Skip to content

Schedule a job

workloads:
nightly-dump:
role: Job
image: postgres:17
command: ["sh", "-c", "pg_dump \"$POSTGRES_URL\" | gzip > /backups/$(date -u +%F).sql.gz"]
dataEffect: None
needs: [postgres]
volumes: [{name: backups, path: /backups}]
schedule:
cron: "0 2 * * *"
timezone: Europe/Berlin
timeout: 45m
catchUp: true

A schedule is translated into a timer on the host. It fires without any Onebox process running and survives a reboot. There is no daemon, no listening port, and nothing to keep alive.

That design decision has a visible consequence elsewhere: encrypted environment entries are decrypted into the release when it is staged and stay there, because a timer firing at 02:00 must resolve the values the deploy resolved.

Each run is a systemd oneshot with a wall-time limit. timeout defaults to 1h; when it expires, systemd terminates the run and records timeout as its result. Set a longer duration for jobs that legitimately need it.

catchUp defaults to true: if the host was off at the scheduled time, the timer runs once after it returns. Set it to false for time-sensitive work that should be skipped rather than run late.

By default a failed run is not retried; the next attempt is the next cron elapse, which is the right answer for anything that fires every few minutes. For a job whose failures are usually transient, a rate limit or an upstream that was busy for a moment, declare a bounded retry:

schedule:
cron: "0 * * * *"
timeout: 45m
retry: {attempts: 3, backoff: 30s, max_backoff: 10m}

attempts counts every attempt including the first, so the default of 1 is today’s single run. After a non-zero exit the runner sleeps backoff, doubles it after each further failure, and never sleeps longer than max_backoff. Every attempt runs inside the same timer firing, under the same locks and the same timeout; when the timeout expires the run ends, with no further attempt.

Validation refuses a retry whose worst-case backoff is not smaller than the timeout, because the last attempt could never start and the record would say it did. With deployLock: Exclusive the deploy lock is held through the sleeps, so a deploy waits for the run to finish; with deployLock: Pinned the release lease is held instead and every attempt runs the same release.

Deployment coordination is exclusive by default

Section titled “Deployment coordination is exclusive by default”

An exclusive host-fired job takes an application-wide kernel lock for its whole container run. An ordinary pinned job takes that rendezvous as a shared reader until it establishes its immutable release lease; a durable pinned job retains the reader through checkpoint preparation and publication. Different pinned jobs may still start together. Every operation that establishes Onebox’s fenced application lock takes the rendezvous as an exclusive writer while publishing its ownership. Readers and writers wait briefly for a handoff; the wait is shortened for a job whose own timeout is near the ten-second default. A timer that still collides does not modify Docker beside a deploy: it records a skipped run with the reason and exits cleanly. An application lock older than its TTL is treated as expired, exactly as a deploy treats it, so a runner that died mid-operation cannot silence a timer forever.

The target must provide flock (part of util-linux on supported Linux hosts). Onebox refuses to install or run schedules when that serialization primitive is missing.

For a long-running job that does not change shared data, opt into a pinned release instead:

dataEffect: None
schedule:
cron: "0 2 * * *"
timeout: 6h
deployLock: Pinned

The runner briefly meets deployment under the same scheduling mutex, resolves current to one immutable release, and takes a shared lease on that release. It then releases the application-wide mutex and keeps only a per-job lock, so a data-effect-free deployment can proceed but a second instance of the same job cannot overlap. Deployments containing migration, destructive, or unknown-effect jobs, and deployments containing untyped lifecycle hooks, wait for the pinned job to finish. Other application operations also remain exclusive. Release cleanup preserves the leased Compose document, environment files, and release-bound mounts until the job exits. ob status reports the policy, the pinned release, start time, and timeout while it runs.

This is deliberately fail-closed. deployLock: Pinned is accepted only for a Onebox-rendered job declaring dataEffect: None. Migration, destructive, unknown-effect, and adopted-Compose jobs remain exclusive because pinning their files cannot prove that a concurrent deployment is safe for the data or external files they use. Destroy and every other non-deploy application operation also refuse while any pinned release is leased.

The ob binary is agentless: upgrading it on the operator workstation does not silently connect to or mutate a target. Existing timers therefore keep the unit content written by the previous runner until the next deploy or an explicit schedule apply:

Terminal window
ob schedule apply

This rewrites every declared timer, service, runner, and failure notifier from the current contract, and removes units for jobs no longer declared. It runs under the same application lock, schedule mutex, fence, and journal boundary as a deploy, but it does not stage or activate a release. Run it after upgrading Onebox when an application may not be deployed again soon.

Run records are written by the unit, so a host whose units predate them has none until this command or a deploy rewrites them. Until then ob status falls back to what systemd retained, reports a failure as it always did, and says that no run record exists yet.

When a run ends, for any reason, the unit’s ExecStopPost writes one record to the host journal with the syslog identifier onebox-run and the job’s unit in its ONEBOX_UNIT field, so journalctl SYSLOG_IDENTIFIER=onebox-run ONEBOX_UNIT=onebox-job-<job> -o cat on the host is the raw history:

{"run":"a3f9…","job":"nightly-dump","trigger":"timer","operation":"","release":"20260905-140000-ab12cd","started_at":"2026-09-05T02:00:01Z","finished_at":"2026-09-05T02:04:37Z","duration_s":276,"attempts":2,"exit_status":0,"outcome":"success","reason":"","inputs":{}}

run is systemd’s invocation id, so the record and the run’s own log share a key. outcome is one of success, failure, timeout, or skipped. A skip is a firing that met a running instance of the same job, timed out waiting for the application scheduling rendezvous, or found an application operation holding the deploy lock. The runner writes the reason and exits before any container starts, so the unit is not failed and the job’s own exit status is never mistaken for a skip. Skips are recorded because a job that is silently never running looks exactly like one that works: one skip is timing, and three in a row are reported by ob status as an issue.

A skip is news about timing, not about the job, so it never clears a failure: ob status keeps reporting the newest run that actually happened, and says that nothing has run since. A job that selects skipped in notify is told its run did not happen, not that it failed.

Read the records back from the workstation:

Terminal window
ob schedule list # every job, its timer state and next elapse
ob job history nightly-dump # records, newest first; -n 50 for more
ob job logs nightly-dump # the journal of the newest run
ob job logs nightly-dump --run a3f9…

ob status reads the same records. Each scheduled job’s line carries the next elapse, the last outcome with its duration and attempt count, and, while a run is in progress, the attempt it is on. A last outcome of failure or timeout is reported as divergence:

schedule nightly-dump last run failed: timeout (exit 143) ⚠

Retention is the journal’s. A host whose journal lives in memory keeps records only since its last boot; ob status says so on the job’s line, and the fix is a persistent journal (/var/log/journal), which the supported images have by default.

schedule.notify selects which outcomes send the configured notifications. The default is [failure, timeout], which is what the failure notifier always did. Add success for a job whose completion someone waits for, and skipped for a job whose firings collide often enough that silence would hide it:

schedule:
cron: "0 2 * * *"
notify: [failure, timeout, skipped]

Each selected outcome sends a bounded, fail-open POST to every webhook whose own on list accepts that class of outcome. The payload carries the run id as its deploy_id, so ob job logs <job> --run <id> finds the run; it carries nothing else about the run, because notifications cross the host trust boundary and diagnostics stay on the host. Onebox writes the webhook handler root-only beside the unit, so credentials in webhook paths do not appear in ExecStart or ExecStopPost. A delivery failure is recorded in the unit’s journal and never replaces the job’s own result. The webhook must be reachable from the target host; timer delivery uses its network path, not the operator workstation’s.

0 2 * * * ✓ every day at 02:00
0 2 * * 1 ✓ every Monday at 02:00
0 2 1 * 1 ✗ schedule_untranslatable

The last one declares a day-of-month and a day-of-week. Cron treats that as “either matches” — it fires on the 1st and on every Monday. Onebox refuses it at load rather than running on days nobody chose.

timezone takes an IANA zone name and defaults to UTC. Generated timers use one-second accuracy so systemd does not coalesce distinct cron minutes into its default one-minute wake-up window.

dataEffect is required on every job, scheduled or not. A nightly report is none; a nightly prune is destructive. The rollback and abort gates read it, and a job that lies about it defeats them.

Terminal window
ob schedule pause nightly-dump --reason "upstream is returning garbage"
ob schedule resume nightly-dump

Pausing stops the timer and nothing else. The units stay installed and later deploys keep updating them, so a fix still lands on the host while the job stays stopped. A run already under way is left alone: this stops the next firing, it does not kill work in progress.

The pause is recorded on the host, not in the project, which is what lets it survive a deploy. Reconciliation will not start a paused timer, so a pause cannot be undone by someone else’s release. --reason is required, and is kept with the operator and the time, because a job that is deliberately not running is indistinguishable from one that is broken:

schedule nightly-dump PAUSED — by [email protected]; since 2026-09-06T10:00:00Z; upstream is returning garbage

ob status prints that line whether or not anything else is wrong, and does not count the stopped timer as divergence. It does still report everything else about the job: a unit that will not load, or a run that failed, is printed beside the pause and still fails the command, because a pause explains a stopped timer and nothing more. A stopped timer with no pause behind it is also still a fault. ob schedule list shows a paused job’s timer as paused rather than inactive, so the table does not read the same for a job somebody stopped and one that broke.

Both ob schedule pause and ob schedule resume are journaled and appear in ob audit with the operator. Neither will record a change that did not happen: pausing an already-paused job is refused rather than overwriting who stopped it and why, and resuming a job nobody paused is refused rather than logged. Both take --break-lock, because a running job holds the application lock for its whole run and a crashed deploy leaves one behind.

Resuming starts the timer again; the next run is the next scheduled elapse. It does not run the job now, and does not make up firings missed while paused — ob job run is how you ask for an immediate run.

Deleting a job from the project clears its pause along with its units, so a name reused later does not come back stopped for a reason from a previous life.

A scheduled job may declare inputs: named parameters that reach the container as environment variables. A timer firing uses the defaults; an operator may run the job now and override them.

workloads:
source-sync:
role: Job
image: ghcr.io/acme/ingest:1.8.2
command: ["./ingest", "sync"]
dataEffect: None
inputs:
SOURCE:
enum: [catalog, prices, reviews]
default: catalog
description: Which upstream to sync.
SINCE:
pattern: '^([0-9]{4}-[0-9]{2}-[0-9]{2})?$'
default: ""
description: Only records changed since this date. Empty means the stored cursor.
schedule:
cron: "0 * * * *"
deployLock: Pinned
Terminal window
ob job run source-sync --input SOURCE=prices --input SINCE=2026-09-01

Names are upper-case identifiers outside Onebox’s ONEBOX_ namespace and may not repeat an env key. Each input declares exactly one of enum or pattern, and a default that satisfies it; a pattern matches the whole value. Whatever the pattern allows, a value may not contain a double quote, a backslash, or a control character, and is at most 256 bytes. Those rules are what let the runner hand values to the container without escaping anything.

ob job run validates every value on the workstation, refuses if the unit is already running or an operator run is still pending, journals the request as schedule_run with the operator and the inputs, and starts the unit. It holds the application lock only while writing and journaling, because the runner skips a run that meets that lock. The host record of the run carries the operation id, so ob audit and ob job history join on it. The command follows the host-supervised unit by default; --detach returns once the unit accepts the run.

Only a job with dataEffect: None accepts inputs. Every operator invocation uses a sealed plan; migration jobs additionally retain their approval and backup-report gates and remain attached so Onebox can capture result evidence.

The runner tells an operator activation from a timer firing by the TRIGGER_UNIT variable systemd sets on timer activations, which that project introduced in version 252. Ubuntu 24.04 and Debian 12 qualify; Ubuntu 22.04 and Debian 11 do not.

Scheduled jobs still run on those older hosts, unchanged. Two things narrow:

  • A job that declares inputs is refused, by ob preflight and by ob deploy, before anything is staged. Without TRIGGER_UNIT the next timer firing would read the file meant for an operator’s run.
  • ob job run is refused for the same reason. The timer keeps firing.

The records on such a host say trigger: unknown, because the runner cannot observe what started it and will not guess.

Keep durable checkpoints and resume explicitly

Section titled “Keep durable checkpoints and resume explicitly”

Opt into durable execution when a failed job needs to continue with its original inputs. Onebox saves orchestration state on the managed host; the application keeps responsibility for domain checkpoints and idempotent effects. This requires a native Onebox job with schedule, deploymentPhase: None (the default), and dataEffect: None. Adopted Compose jobs and release-phase jobs are refused.

For one command, add execution: {retention: 168h} to the job. Onebox treats the job’s existing command as one step with ID main. An ingestion job that already keeps its own domain checkpoints can use this form and keep them where they are.

For several ordered commands, declare steps:

workloads:
refresh:
role: Job
image: ghcr.io/acme/catalog:1.8.2
dataEffect: None
schedule:
cron: "0 * * * *"
timeout: 45m
retry: {attempts: 2, backoff: 30s, max_backoff: 10m}
execution:
retention: 168h
steps:
- id: sync
command: ["./catalog", "sync"]
outputs: [RELEASE]
- id: index
command: ["./catalog", "index"]
inputs: {RELEASE_ID: "sync.RELEASE"}
retry: {attempts: 3}

Each step uses the job’s image and entrypoint. command is an argument list, passed without shell evaluation. In this example, sync writes the application release it produced, and index reads it through the RELEASE_ID environment variable. That value is separate from Onebox’s deployment release ID.

The application writes a JSON object to the path in ONEBOX_OUTPUT_FILE, for example {"RELEASE":"catalog-20260906"}. The object must contain exactly the declared output keys, and every value must be a string without NUL. Each value is at most 4096 UTF-8 bytes; the output file and its normalized JSON encoding must each fit within 16384 bytes. Keep large exports elsewhere and save a small reference or checksum. Inputs and outputs are persisted operational metadata: never put credentials or other secrets in them. Current credentials are resolved for an attempt; the execution record does not snapshot them.

These bounds intentionally limit step outputs to small operational metadata, such as identifiers, cursors, or checksums. They do not limit how much data a job can process. Thousands of row IDs can exceed the bounds; raising the limits alone would not solve large payload handoff, because step inputs are delivered through environment variables and outputs are retained in execution checkpoints.

For a large result set, the application stores the data in its own database or artifact storage and returns a small reference:

steps:
- id: sync
command: ["./ingest", "sync"]
# `sync` stores the batch and writes {"BATCH":"batch-id"}
# to ONEBOX_OUTPUT_FILE.
outputs: [BATCH]
- id: index
command: ["./ingest", "index"]
inputs: {BATCH_ID: "sync.BATCH"}

index reads the batch using the BATCH_ID environment variable. To resume against the same data, the application must save the batch durably before returning its reference and retain it unchanged for as long as the execution remains resumable. Onebox retains the reference; the application owns the batch’s availability, access control, and cleanup. References must not contain secrets, such as credentials embedded in a URL.

A single command is also appropriate when the application already owns recovery for the whole operation. It can keep its own checkpoints and reconcile pending work without a separate batch handoff. Retrying that command starts it again, so its recovery logic must handle completed effects and unfinished work safely. Use separate steps when independent retries and resume boundaries are useful.

Step IDs are unique. An input reference names a declared output of an earlier step as stepID.OUTPUT; forward and self references are refused. Input and output names are upper-case identifiers outside ONEBOX_; step input names must not collide with the job’s original inputs or literal env keys. There may be at most 32 steps, 32 original inputs, and 32 inputs and outputs per step.

A step succeeds only after its exit status and validated outputs have been saved together. A failed step retries independently, inheriting schedule.retry unless it declares its own retry block. The sum of all steps’ worst-case backoff must fit below schedule.timeout.

Terminal window
ob job run refresh # start a new execution
ob execution list # newest executions; -n 50 for more
ob execution inspect <id> # original inputs, steps, attempts, eligibility
ob execution resume <id> --wait # skip completed steps; use saved outputs
ob execution abandon <id> # end resumability and release retention hold

Resume cannot change inputs. To use different inputs, start a new execution with ob job run --input NAME=VALUE. Timer firings also create new executions; they do not resume failed work automatically. After a reboot, inspect interrupted work and request resume explicitly. A pending failed execution does not stop future timer firings; use ob schedule pause <job> --reason "..." if needed.

Inspection includes each activation’s systemd invocation ID. Read its logs with ob job logs refresh --run <invocation-id>. The execution ID stays the same across activations; an invocation ID identifies one activation’s journal output. Run logs keep the journal’s retention independently of saved execution state.

The job receives ONEBOX_EXECUTION_ID and ONEBOX_STEP_ID, both stable across retries and resumes, plus a fresh ONEBOX_ATTEMPT_ID for each attempt. Use the stable IDs to deduplicate effects in the application. A crash can occur after an external effect succeeds but before Onebox saves completion, so that step can run again. Durable checkpoints provide at-least-once attempts, not an exactly-once guarantee for application effects.

retention defaults to 168h and may be at most 30d. It sets the resume window from execution creation; resuming does not extend it. An execution can have at most 100 activations, including its first run. Successful and abandoned executions cannot resume. Expired inactive executions stop protecting their release; running or uncertain work is protected conservatively until it can be safely abandoned. Abandonment refuses active work and keeps the saved evidence available for inspection until its retention deadline. Starting a new execution prunes expired records that are not still recorded as running; interrupted records left running by a crash require explicit abandonment.

schedule.timeout bounds each activation, including every step and retry sleep. A manual resume starts a new activation with that same bound. Deployment locks, overlap protection, pinned-release behavior, and notifications cover the whole activation. A failed execution awaiting manual resume holds its durable release reference, not a deployment lock.

Resume requires the original release to still be current, the installed workflow definition to match, and the original image to be locally available with the same identity. Onebox also checks release files, managed-service container identities, images and start times, and its recorded data-change generation. Data-changing jobs and arbitrary ob exec commands invalidate that generation before running, because Onebox cannot prove an arbitrary command is read-only. A service restart or host reboot can therefore prevent resume even when release files survive. These checks cannot establish compatibility after arbitrary out-of-band schema changes; they do not replace application-level compatibility checks. There is no cross-release adaptation.

The managed host needs Python 3.8 or later at /usr/bin/python3 in addition to the scheduler’s existing systemd and flock requirements. Onebox installs a small standard-library helper invoked only for executions and inspection; it does not run a resident worker service. Durable opt-in also requires the systemd 252 trigger support described above. State lives under the application’s schedule/executions directory as protected JSON checkpoints and survives a runner or host restart when the host disk survives. It is not replicated to another host.

Terminal window
ob job plan nightly-dump --out ob-job-plan.json
ob approve --plan ob-job-plan.json --out ob-job-approval.json
ob job run --plan ob-job-plan.json --approval ob-job-approval.json

When the operator job also declares schedule, Onebox submits this run to the same installed systemd unit used by its timer. The host therefore owns the container, timeout, overlap lock, retries, and run record if the SSH session or operator terminal disappears. The command follows the unit by default and shows elapsed time. Ctrl-C stops following but does not stop the host job; use --detach to return immediately after the unit accepts it. Inspect either form with ob job history <job> and ob job logs <job>.

Jobs without schedule keep the direct foreground path. Migration jobs also stay attached because their result evidence is part of the approved operation.

ob job plan accepts a job when operatorRun is allowed. This property is independent of deployment: deploymentPhase decides whether the job runs in the deploy graph, while operatorRun decides whether an operator may invoke it. By default, phase none allows operator runs; release-phase jobs disable them unless explicitly enabled.

A declared deploymentPhase: None job remains in the digest-pinned release runtime but never joins the deploy graph. Its job plan binds the current serving release, runtime digest, immutable image, data effect, target, and expiry. Automation supplies the saved plan and its separately recorded local confirmation; migration jobs may also need the exact plan-bound backup report.

For an undeclared emergency command, the escape hatch remains:

Terminal window
ob exec --reason "incident investigation" <workload|service> -- <command>

For a job that participates automatically in a deploy, choose a release phase:

deploymentPhase: PreRelease
operatorRun: allowed # optional: also permit ob job plan/run

The three properties are independent. deploymentPhase controls deploy hooks, schedule controls timer activation, and operatorRun controls explicit invocation. A job declaring schedule: and deploymentPhase: PreRelease runs both at its cron time and on every deploy. For a timer-only job, leave deploymentPhase at none; set operatorRun: Disabled too if operators must not start it explicitly.

See workloads for every field.