Schedule a job
workloads: nightly-dump: role: Job image: postgres:17 command: ["sh", "-c", "pg_dump \"$POSTGRES_URL\" | gzip > /backups/$(date -u +%F).sql.gz"] dataEffect: None needs: [postgres] volumes: [{name: backups, path: /backups}] schedule: cron: "0 2 * * *" timezone: Europe/Berlin timeout: 45m catchUp: trueIt becomes a host timer
Section titled “It becomes a host timer”A schedule is translated into a timer on the host. It fires without any
Onebox process running and survives a reboot. There is no daemon, no listening
port, and nothing to keep alive.
That design decision has a visible consequence elsewhere: encrypted environment entries are decrypted into the release when it is staged and stay there, because a timer firing at 02:00 must resolve the values the deploy resolved.
Each run is a systemd oneshot with a wall-time limit. timeout defaults to
1h; when it expires, systemd terminates the run and records timeout as its
result. Set a longer duration for jobs that legitimately need it.
catchUp defaults to true: if the host was off at the scheduled time, the
timer runs once after it returns. Set it to false for time-sensitive work that
should be skipped rather than run late.
Retry inside one firing
Section titled “Retry inside one firing”By default a failed run is not retried; the next attempt is the next cron elapse, which is the right answer for anything that fires every few minutes. For a job whose failures are usually transient, a rate limit or an upstream that was busy for a moment, declare a bounded retry:
schedule: cron: "0 * * * *" timeout: 45m retry: {attempts: 3, backoff: 30s, max_backoff: 10m}attempts counts every attempt including the first, so the default of 1 is
today’s single run. After a non-zero exit the runner sleeps backoff, doubles
it after each further failure, and never sleeps longer than max_backoff. Every
attempt runs inside the same timer firing, under the same locks and the same
timeout; when the timeout expires the run ends, with no further attempt.
Validation refuses a retry whose worst-case backoff is not smaller than the
timeout, because the last attempt could never start and the record would say
it did. With deployLock: Exclusive the deploy lock is held through the
sleeps, so a deploy waits for the run to finish; with deployLock: Pinned the
release lease is held instead and every attempt runs the same release.
Deployment coordination is exclusive by default
Section titled “Deployment coordination is exclusive by default”An exclusive host-fired job takes an application-wide kernel lock for its whole
container run. An ordinary pinned job takes that rendezvous as a shared reader
until it establishes its immutable release lease; a durable pinned job retains
the reader through checkpoint preparation and publication. Different pinned
jobs may still start together. Every operation that establishes Onebox’s fenced
application lock takes the rendezvous as an exclusive writer while publishing
its ownership. Readers and writers wait briefly for a handoff; the wait is
shortened for a job whose own timeout is near the ten-second default. A timer
that still collides does not modify Docker beside a deploy: it records a
skipped run with the reason and exits cleanly. An application lock older than
its TTL is treated as expired, exactly as a deploy treats it, so a runner that
died mid-operation cannot silence a timer forever.
The target must provide flock (part of util-linux on supported Linux hosts).
Onebox refuses to install or run schedules when that serialization primitive is
missing.
For a long-running job that does not change shared data, opt into a pinned release instead:
dataEffect: Noneschedule: cron: "0 2 * * *" timeout: 6h deployLock: PinnedThe runner briefly meets deployment under the same scheduling mutex, resolves
current to one immutable release, and takes a shared lease on that release.
It then releases the application-wide mutex and keeps only a per-job lock, so a
data-effect-free deployment can proceed but a second instance of the same job
cannot overlap. Deployments containing migration, destructive, or unknown-effect
jobs, and deployments containing untyped lifecycle hooks, wait for the pinned
job to finish. Other application operations also remain exclusive. Release
cleanup preserves the leased Compose document, environment files, and
release-bound mounts until the job exits. ob status reports the policy, the
pinned release, start time, and timeout while it runs.
This is deliberately fail-closed. deployLock: Pinned is accepted only for a
Onebox-rendered job declaring dataEffect: None. Migration, destructive,
unknown-effect, and adopted-Compose jobs remain exclusive because pinning their
files cannot prove that a concurrent deployment is safe for the data or
external files they use. Destroy and every other non-deploy application
operation also refuse while any pinned release is leased.
Reconcile after upgrading Onebox
Section titled “Reconcile after upgrading Onebox”The ob binary is agentless: upgrading it on the operator workstation does not
silently connect to or mutate a target. Existing timers therefore keep the unit
content written by the previous runner until the next deploy or an explicit
schedule apply:
ob schedule applyThis rewrites every declared timer, service, runner, and failure notifier from the current contract, and removes units for jobs no longer declared. It runs under the same application lock, schedule mutex, fence, and journal boundary as a deploy, but it does not stage or activate a release. Run it after upgrading Onebox when an application may not be deployed again soon.
Run records are written by the unit, so a host whose units predate them has
none until this command or a deploy rewrites them. Until then ob status falls
back to what systemd retained, reports a failure as it always did, and says
that no run record exists yet.
Every run leaves a record
Section titled “Every run leaves a record”When a run ends, for any reason, the unit’s ExecStopPost writes one record
to the host journal with the syslog identifier onebox-run and the job’s unit in
its ONEBOX_UNIT field, so journalctl SYSLOG_IDENTIFIER=onebox-run ONEBOX_UNIT=onebox-job-<job> -o cat on the host is the raw history:
{"run":"a3f9…","job":"nightly-dump","trigger":"timer","operation":"","release":"20260905-140000-ab12cd","started_at":"2026-09-05T02:00:01Z","finished_at":"2026-09-05T02:04:37Z","duration_s":276,"attempts":2,"exit_status":0,"outcome":"success","reason":"","inputs":{}}run is systemd’s invocation id, so the record and the run’s own log share a
key. outcome is one of success, failure, timeout, or skipped. A skip
is a firing that met a running instance of the same job, timed out waiting for
the application scheduling rendezvous, or found an application operation
holding the deploy lock. The runner writes the reason and exits before any
container starts, so the unit is not failed and the job’s own exit status is
never mistaken for a skip. Skips are recorded because a job that is silently
never running looks exactly like one that works: one skip is timing, and three
in a row are reported by ob status as an issue.
A skip is news about timing, not about the job, so it never clears a failure:
ob status keeps reporting the newest run that actually happened, and says
that nothing has run since. A job that selects skipped in notify is told
its run did not happen, not that it failed.
Read the records back from the workstation:
ob schedule list # every job, its timer state and next elapseob job history nightly-dump # records, newest first; -n 50 for moreob job logs nightly-dump # the journal of the newest runob job logs nightly-dump --run a3f9…ob status reads the same records. Each scheduled job’s line carries the next
elapse, the last outcome with its duration and attempt count, and, while a run
is in progress, the attempt it is on. A last outcome of failure or timeout
is reported as divergence:
schedule nightly-dump last run failed: timeout (exit 143) ⚠Retention is the journal’s. A host whose journal lives in memory keeps records
only since its last boot; ob status says so on the job’s line, and the fix is
a persistent journal (/var/log/journal), which the supported images have by
default.
Notifications per outcome
Section titled “Notifications per outcome”schedule.notify selects which outcomes send the configured notifications. The
default is [failure, timeout], which is what the failure notifier always did.
Add success for a job whose completion someone waits for, and skipped for a
job whose firings collide often enough that silence would hide it:
schedule: cron: "0 2 * * *" notify: [failure, timeout, skipped]Each selected outcome sends a bounded, fail-open POST to every webhook whose
own on list accepts that class of outcome. The payload carries the run id as
its deploy_id, so ob job logs <job> --run <id> finds the run; it
carries nothing else about the run, because notifications cross the host trust
boundary and diagnostics stay on the host. Onebox writes the webhook handler
root-only beside the unit, so credentials in webhook paths do not appear in
ExecStart or ExecStopPost. A delivery failure is recorded in the unit’s
journal and never replaces the job’s own result. The webhook must be reachable
from the target host; timer delivery uses its network path, not the operator
workstation’s.
Cron is translated exactly, or refused
Section titled “Cron is translated exactly, or refused”0 2 * * * ✓ every day at 02:000 2 * * 1 ✓ every Monday at 02:000 2 1 * 1 ✗ schedule_untranslatableThe last one declares a day-of-month and a day-of-week. Cron treats that as “either matches” — it fires on the 1st and on every Monday. Onebox refuses it at load rather than running on days nobody chose.
timezone takes an IANA zone name and defaults to UTC.
Generated timers use one-second accuracy so systemd does not coalesce distinct
cron minutes into its default one-minute wake-up window.
Jobs still declare a data effect
Section titled “Jobs still declare a data effect”dataEffect is required on every job, scheduled or not. A nightly report is
none; a nightly prune is destructive. The rollback and abort gates read it,
and a job that lies about it defeats them.
Stopping a job without deleting it
Section titled “Stopping a job without deleting it”ob schedule pause nightly-dump --reason "upstream is returning garbage"ob schedule resume nightly-dumpPausing stops the timer and nothing else. The units stay installed and later deploys keep updating them, so a fix still lands on the host while the job stays stopped. A run already under way is left alone: this stops the next firing, it does not kill work in progress.
The pause is recorded on the host, not in the project, which is what lets it
survive a deploy. Reconciliation will not start a paused timer, so a pause
cannot be undone by someone else’s release. --reason is required, and is
kept with the operator and the time, because a job that is deliberately not
running is indistinguishable from one that is broken:
schedule nightly-dump PAUSED — by [email protected]; since 2026-09-06T10:00:00Z; upstream is returning garbageob status prints that line whether or not anything else is wrong, and does
not count the stopped timer as divergence. It does still report everything
else about the job: a unit that will not load, or a run that failed, is
printed beside the pause and still fails the command, because a pause explains
a stopped timer and nothing more. A stopped timer with no pause behind it is
also still a fault. ob schedule list shows a paused job’s timer as paused
rather than inactive, so the table does not read the same for a job somebody
stopped and one that broke.
Both ob schedule pause and ob schedule resume are journaled and appear in
ob audit with the operator. Neither will record a change that did not
happen: pausing an already-paused job is refused rather than overwriting who
stopped it and why, and resuming a job nobody paused is refused rather than
logged. Both take --break-lock, because a running job holds the application
lock for its whole run and a crashed deploy leaves one behind.
Resuming starts the timer again; the next run is the next scheduled elapse. It
does not run the job now, and does not make up firings missed while paused —
ob job run is how you ask for an immediate run.
Deleting a job from the project clears its pause along with its units, so a name reused later does not come back stopped for a reason from a previous life.
Run one now with inputs
Section titled “Run one now with inputs”A scheduled job may declare inputs: named parameters that reach the container as environment variables. A timer firing uses the defaults; an operator may run the job now and override them.
workloads: source-sync: role: Job image: ghcr.io/acme/ingest:1.8.2 command: ["./ingest", "sync"] dataEffect: None inputs: SOURCE: enum: [catalog, prices, reviews] default: catalog description: Which upstream to sync. SINCE: pattern: '^([0-9]{4}-[0-9]{2}-[0-9]{2})?$' default: "" description: Only records changed since this date. Empty means the stored cursor. schedule: cron: "0 * * * *" deployLock: Pinnedob job run source-sync --input SOURCE=prices --input SINCE=2026-09-01Names are upper-case identifiers outside Onebox’s ONEBOX_ namespace and may
not repeat an env key. Each input declares exactly one of enum or
pattern, and a default that satisfies it; a pattern matches the whole
value. Whatever the pattern allows, a value may not contain a double quote, a
backslash, or a control character, and is at most 256 bytes. Those rules are
what let the runner hand values to the container without escaping anything.
ob job run validates every value on the workstation, refuses if the
unit is already running or an operator run is still pending, journals the request
as schedule_run with the operator and the inputs, and starts the unit. It
holds the application lock only while writing and journaling, because the
runner skips a run that meets that lock. The host record of the run carries the
operation id, so ob audit and ob job history join on it. The command follows
the host-supervised unit by default; --detach returns once the unit accepts
the run.
Only a job with dataEffect: None accepts inputs. Every operator invocation
uses a sealed plan; migration jobs additionally retain their approval and
backup-report gates and remain attached so Onebox can capture result evidence.
The runner tells an operator activation from a timer firing by the TRIGGER_UNIT
variable systemd sets on timer activations, which that project introduced in
version 252. Ubuntu 24.04 and Debian 12 qualify; Ubuntu 22.04 and Debian 11 do
not.
Scheduled jobs still run on those older hosts, unchanged. Two things narrow:
- A job that declares
inputsis refused, byob preflightand byob deploy, before anything is staged. WithoutTRIGGER_UNITthe next timer firing would read the file meant for an operator’s run. ob job runis refused for the same reason. The timer keeps firing.
The records on such a host say trigger: unknown, because the runner cannot
observe what started it and will not guess.
Keep durable checkpoints and resume explicitly
Section titled “Keep durable checkpoints and resume explicitly”Opt into durable execution when a failed job needs to continue with its original
inputs. Onebox saves orchestration state on the managed host; the application
keeps responsibility for domain checkpoints and idempotent effects. This requires
a native Onebox job with schedule, deploymentPhase: None (the default), and
dataEffect: None. Adopted Compose jobs and release-phase jobs are refused.
For one command, add execution: {retention: 168h} to the job. Onebox treats the
job’s existing command as one step with ID main. An ingestion job that already
keeps its own domain checkpoints can use this form and keep them where they are.
For several ordered commands, declare steps:
workloads: refresh: role: Job image: ghcr.io/acme/catalog:1.8.2 dataEffect: None schedule: cron: "0 * * * *" timeout: 45m retry: {attempts: 2, backoff: 30s, max_backoff: 10m} execution: retention: 168h steps: - id: sync command: ["./catalog", "sync"] outputs: [RELEASE] - id: index command: ["./catalog", "index"] inputs: {RELEASE_ID: "sync.RELEASE"} retry: {attempts: 3}Each step uses the job’s image and entrypoint. command is an argument list,
passed without shell evaluation. In this example, sync writes the application
release it produced, and index reads it through the RELEASE_ID environment
variable. That value is separate from Onebox’s deployment release ID.
The application writes a JSON object to the path in ONEBOX_OUTPUT_FILE, for
example {"RELEASE":"catalog-20260906"}. The object must contain exactly the
declared output keys, and every value must be a string without NUL. Each value
is at most 4096 UTF-8 bytes; the output file and its normalized JSON encoding
must each fit within 16384 bytes. Keep large exports elsewhere and save a small
reference or checksum. Inputs and outputs are persisted operational metadata:
never put credentials or other secrets in them. Current credentials are
resolved for an attempt; the execution record does not snapshot them.
Pass large results by reference
Section titled “Pass large results by reference”These bounds intentionally limit step outputs to small operational metadata, such as identifiers, cursors, or checksums. They do not limit how much data a job can process. Thousands of row IDs can exceed the bounds; raising the limits alone would not solve large payload handoff, because step inputs are delivered through environment variables and outputs are retained in execution checkpoints.
For a large result set, the application stores the data in its own database or artifact storage and returns a small reference:
steps: - id: sync command: ["./ingest", "sync"] # `sync` stores the batch and writes {"BATCH":"batch-id"} # to ONEBOX_OUTPUT_FILE. outputs: [BATCH] - id: index command: ["./ingest", "index"] inputs: {BATCH_ID: "sync.BATCH"}index reads the batch using the BATCH_ID environment variable. To resume
against the same data, the application must save the batch durably before
returning its reference and retain it unchanged for as long as the execution
remains resumable. Onebox retains the reference; the application owns the batch’s
availability, access control, and cleanup. References must not contain secrets,
such as credentials embedded in a URL.
A single command is also appropriate when the application already owns recovery for the whole operation. It can keep its own checkpoints and reconcile pending work without a separate batch handoff. Retrying that command starts it again, so its recovery logic must handle completed effects and unfinished work safely. Use separate steps when independent retries and resume boundaries are useful.
Step IDs are unique. An input reference names a declared output of an earlier
step as stepID.OUTPUT; forward and self references are refused. Input and output
names are upper-case identifiers outside ONEBOX_; step input names must not
collide with the job’s original inputs or literal env keys. There may be at
most 32 steps, 32 original inputs, and 32 inputs and outputs per step.
A step succeeds only after its exit status and validated outputs have been
saved together. A failed step retries independently, inheriting schedule.retry
unless it declares its own retry block. The sum of all steps’ worst-case backoff
must fit below schedule.timeout.
Inspect and recover
Section titled “Inspect and recover”ob job run refresh # start a new executionob execution list # newest executions; -n 50 for moreob execution inspect <id> # original inputs, steps, attempts, eligibilityob execution resume <id> --wait # skip completed steps; use saved outputsob execution abandon <id> # end resumability and release retention holdResume cannot change inputs. To use different inputs, start a new execution with
ob job run --input NAME=VALUE. Timer firings also create new executions;
they do not resume failed work automatically. After a reboot, inspect interrupted
work and request resume explicitly. A pending failed execution does not stop
future timer firings; use ob schedule pause <job> --reason "..." if needed.
Inspection includes each activation’s systemd invocation ID. Read its logs with
ob job logs refresh --run <invocation-id>. The execution ID stays the same
across activations; an invocation ID identifies one activation’s journal output.
Run logs keep the journal’s retention independently of saved execution state.
The job receives ONEBOX_EXECUTION_ID and ONEBOX_STEP_ID, both stable across
retries and resumes, plus a fresh ONEBOX_ATTEMPT_ID for each attempt. Use the
stable IDs to deduplicate effects in the application. A crash can occur after
an external effect succeeds but before Onebox saves completion, so that step
can run again. Durable checkpoints provide at-least-once attempts, not an
exactly-once guarantee for application effects.
Bounds and compatibility
Section titled “Bounds and compatibility”retention defaults to 168h and may be at most 30d. It sets the resume window
from execution creation; resuming does not extend it. An execution can have at
most 100 activations, including its first run. Successful and abandoned
executions cannot resume. Expired inactive executions stop protecting their
release; running or uncertain work is protected conservatively until it can be
safely abandoned. Abandonment refuses active work and keeps the saved evidence
available for inspection until its retention deadline. Starting a new execution
prunes expired records that are not still recorded as running; interrupted
records left running by a crash require explicit abandonment.
schedule.timeout bounds each activation, including every step and retry sleep.
A manual resume starts a new activation with that same bound. Deployment locks,
overlap protection, pinned-release behavior, and notifications cover the whole
activation. A failed execution awaiting manual resume holds its durable release
reference, not a deployment lock.
Resume requires the original release to still be current, the installed workflow
definition to match, and the original image to be locally available with the
same identity. Onebox also checks release files, managed-service container
identities, images and start times, and its recorded data-change generation.
Data-changing jobs and arbitrary ob exec commands invalidate that generation
before running, because Onebox cannot prove an arbitrary command is read-only.
A service restart or host reboot can therefore prevent resume even when release
files survive. These checks cannot establish compatibility after arbitrary
out-of-band schema changes; they do not replace application-level compatibility
checks. There is no cross-release adaptation.
The managed host needs Python 3.8 or later at /usr/bin/python3 in addition to
the scheduler’s existing systemd and flock requirements. Onebox installs a
small standard-library helper invoked only for executions and inspection; it
does not run a resident worker service. Durable opt-in also requires the
systemd 252 trigger support described above. State lives under the application’s
schedule/executions directory as protected JSON checkpoints and survives a
runner or host restart when the host disk survives. It is not replicated to
another host.
Running one by hand
Section titled “Running one by hand”ob job plan nightly-dump --out ob-job-plan.jsonob approve --plan ob-job-plan.json --out ob-job-approval.jsonob job run --plan ob-job-plan.json --approval ob-job-approval.jsonWhen the operator job also declares schedule, Onebox submits this run to the
same installed systemd unit used by its timer. The host therefore owns the
container, timeout, overlap lock, retries, and run record if the SSH session or
operator terminal disappears. The command follows the unit by default and
shows elapsed time. Ctrl-C stops following but does not stop the host job; use
--detach to return immediately after the unit accepts it. Inspect either form
with ob job history <job> and ob job logs <job>.
Jobs without schedule keep the direct foreground path. Migration jobs also
stay attached because their result evidence is part of the approved operation.
ob job plan accepts a job when operatorRun is allowed. This property is
independent of deployment: deploymentPhase decides whether the job runs in
the deploy graph, while operatorRun decides whether an operator may invoke
it. By default, phase none allows operator runs; release-phase jobs disable
them unless explicitly enabled.
A declared deploymentPhase: None job remains in the digest-pinned release runtime but
never joins the deploy graph. Its job plan binds the current serving release,
runtime digest, immutable image, data effect, target, and expiry. Automation
supplies the saved plan and its separately recorded local confirmation;
migration jobs may also need the exact plan-bound backup report.
For an undeclared emergency command, the escape hatch remains:
ob exec --reason "incident investigation" <workload|service> -- <command>For a job that participates automatically in a deploy, choose a release phase:
deploymentPhase: PreReleaseoperatorRun: allowed # optional: also permit ob job plan/runThe three properties are independent. deploymentPhase controls deploy hooks,
schedule controls timer activation, and operatorRun controls explicit
invocation. A job declaring schedule: and deploymentPhase: PreRelease
runs both at its cron time and on every deploy. For a timer-only job, leave
deploymentPhase at none; set operatorRun: Disabled too if operators
must not start it explicitly.
See workloads for every field.