API stability
This page states what counts as django-ox's public API, how versions
change, and which Python and Django versions are supported. It is a
promise about compatibility, so you can pin django-ox with confidence.
Public API
These are the supported surfaces. Changes to them are versioned and
announced in the changelog. This page, not a module's
__all__, is the statement of what is public. A module's __all__ may
export names this page does not list; those names are not public.
- The backend path
django_ox.backend.OxBackend, referenced as a string in theTASKSsetting,QUEUESbesideOPTIONS, and everyOPTIONSkey it reads:MAX_ATTEMPTS,LOCK_TIMEOUT,BACKOFF_INITIAL,BACKOFF_MAX,TASK_TIMEOUT,TASK_TIMEOUTS,TASK_TIMEOUT_GRACE,SCHEDULES(with its documented per-schedule keys), andWORKER_CLASS, the dotted path of theWorkersubclassox_workerruns. - The management commands and their flags:
ox_worker,ox_prune,ox_health.ox_worker's exit codes: 0 after a drain, 130 on a forced exit, 75 when the worker recycles itself after a stuck task thread. A worker that finishes under--batchor--max-tasksdrains and exits 0, even when task attempts or individual schedule dispatches failed. Task outcomes stay on their task rows; schedule-scoped failures are reported in logs and leave no tick or task row. Exit 0 means the batch finished, not that every schedule enqueued. An abandoned dispatch pass prevents normal batch-empty completion until a later pass completes; reaching--max-tasksstill ends the run. An invalid--max-tasks, or either flag with--processesabove 1, exits 1 before any worker starts. Under--processes, the supervisor exits 0 when every worker drained or recycled, 1 when a slot hit the restart cap, and otherwise with the first other non-zero worker code. A worker killed by a signal reports128 + the signal number, following the shell convention. - The stored-schedule write functions
django_ox.stored.create_schedule,django_ox.stored.create_schedules,django_ox.stored.update_scheduleanddjango_ox.stored.delete_schedule. - The heartbeat-file protocol: with
ox_worker, one process writesPATH; above one process, the supervisor writesPATH.supervisorand slot i writesPATH.i. The modification time is the signal; file contents are not read or written. Every expected file must be regular and have an age between zero and the configured maximum, inclusive. A passing check means the expected controlling loops have advanced recently, not that tasks are progressing. A launcher that constructs one Worker withheartbeat_fileand then forks it leaves the children touching the same inherited path under their own worker ids. A fresh shared path proves that one writer is alive, not that every child is alive. The documented file-mode JSON fields are public too. - The system check IDs, including the
django_ox.E0xxanddjango_ox.W0xxidentifiers, which you may list inSILENCED_SYSTEM_CHECKS. The IDs are stable; the messages are not. ox_health's exit codes: 0 when every enabled check passes, 1 when a check fails. In file mode, passing means every expected heartbeat file is fresh. A value the command itself rejects exits 1 as well, such as--max-age 0,--max-heartbeat-age 0,--processes 0, a refused flag combination, or a--databasealias that isn't inDATABASES. A value argparse rejects exits 2, such as--max-age nonsense,--max-heartbeat-age nan,--processes 1.5or--format bogus, and so does an unknown flag.- The claim filter hooks
Worker.claim_filter_q()andWorker.claim_filter_sql(), and where their result is applied: the fragment is conjoined to the conditions the candidate select filters on, ahead of its ordering and its limit. The statement around it is not promised. - The timeout helpers
django_ox.deadline()anddjango_ox.remaining(), callable from inside a task.deadline()answers with a wall-clock time, andremaining()with seconds measured on the same monotonic clock the worker enforces the deadline with, so the two can differ by the size of a clock correction.remaining()is the one the watchdog agrees with. - The metrics module
django_ox.stats:queue_stats,ready_count,oldest_ready_age,throughput,failure_rate,last_claim_age,waiting_counts, theQueueStatsdataclass, andDEFAULT_WINDOW, the trailing window the rate functions default to. - The Prometheus surface:
django_ox.metrics.render_prometheus,render_openmetricsandcollector, the viewdjango_ox.views.metricswith theusingargument it takes from the URLconf, thedjango_ox.urlsmodule with itsmetricsroute name, and the metric names and label names listed on the Monitoring page.METRIC_NAMESis that list in code;CONTENT_TYPE_PROMETHEUSandCONTENT_TYPE_OPENMETRICSare the content types the view serves. A scraped name is a contract with every dashboard that reads it, so a rename is a breaking change. Help text is not part of the contract. - The actions module
django_ox.actions:retry,discardandexpire_lease, their accepted states, and their return values;retry_manyanddiscard_many, the selections they accept and their(changed, skipped)return.RETRYABLE_STATUSESandDISCARDABLE_STATUSESare those accepted states in code;UPDATE_CHUNK_SIZEis exported for reading and its value may change. The admin page that calls them is a convenience over this module; its layout is not a contract, the two action names are. - The bulk module
django_ox.bulk:enqueue_many(task, calls), its(args, kwargs)call shape, the input-order return and the all-or-nothing write.INSERT_CHUNK_SIZEis exported for reading; its value may change. - The exceptions
django_ox.exceptions.TaskAbandoned, recorded against tasks whose worker stopped reporting with no attempts left (it records the lost lease, not a cause of failure), anddjango_ox.exceptions.TaskTimeout, aTimeoutErrorraised inside a sync task at its execution deadline and recorded against a timed-out attempt. The timeout comes from the task's own declaration, then its queue'sTASK_TIMEOUTSentry, thenTASK_TIMEOUT. Async tasks are cancelled and their timeout is recorded withTaskTimeout. - The structured-log contract: the event names and stable
extrakeys documented on the Monitoring page. This includes the policy events and keys, even though the declaration API is provisional. It also includesworker_claim_released,worker_claim_recovery_failed,worker_claim_recovery_expiredandworker_claim_release_refused, and their documented keys. These are stable, not provisional. The recovery keys includeclaim_recovery,old_epoch,new_epoch,refunded_attempts,attemptsandworker_ids. Onworker_poll_failed,claim_recoveryis"pending"while a recovery read is owed, ornullotherwise; with psycopg2 it is alwaysnull. Onworker_claim_recovery_failedandworker_claim_recovery_expired, it is always"expired". - The testing helpers
django_ox.testing.ImmediateBackend,django_ox.testing.DummyBackendanddjango_ox.testing.run_tasks. The backends accept policy declarations but do not enforce retries, backoff or timeouts.run_tasks()drains due queued tasks through the configured worker class, including retries and backoff. It does not enforce timeouts.run_tasks()is public and provisional. - The database schema of
OxTaskandOxScheduleTick, evolved only through shipped migrations. django_ox.__version__.
The producer-side API uses the Tasks framework's @task, .enqueue() and
get_result(). Bare tasks follow that framework's contract. django-ox also
accepts the provisional policy declarations below.
Provisional task policy
django_ox.tasks.PolicyTask and django_ox.tasks.BackoffCallback, also
available as django_ox.PolicyTask and django_ox.BackoffCallback, are
public, provisional APIs. This includes the max_attempts, backoff and
timeout fields accepted through @task by OxBackend, and their
representation on result.task.
This surface follows the names and two-argument callback in Django's open
new-features proposals #142 and #144. It makes no promise of compatibility
with whatever Django core eventually ships. timedelta callback returns
are supported; timeout declarations use integer seconds. Retry scheduling
updates run_after, rather than preserving its original value as proposal
142 suggests.
See Configuration for validation, framework support and precedence.
Not public
Everything else is an implementation detail and may change in any release
without notice. That covers the django_ox.worker.Worker internals, the cron
parser (django_ox.cron), the row-to-dataclass conversion
(django_ox.results), the schedule loader (django_ox.schedules), the
supervisor behind --processes (django_ox.supervisor) and the hidden
--worker-index flag it starts each child with, and any name starting
with an underscore. The exact SQL a claim emits and the
model's non-schema helper methods are not part of the contract.
django_ox.heartbeat is also an implementation detail, not a public Python API. The documented heartbeat filenames, modification-time meaning, freshness rule and ox_health JSON fields are public contracts.
In django_ox.tasks, only PolicyTask and BackoffCallback are public,
with the provisional status above. MAX_ATTEMPTS_LIMIT and
validate_policy are implementation details. In django_ox.testing, only
ImmediateBackend, DummyBackend and run_tasks are public.
run_tasks is provisional. django_ox._run_tasks and Worker._task_body
are private implementation details. Timeout implementation classes and
methods, including TaskTimeouts.enabled and TaskTimeouts.for_attempt(),
are not public.
Worker implementation details
Worker internals remain Not public. This includes _handed_off,
_unsettled, _fence, _claims_in_flight, _claim_generation,
_claiming, _claim, _claim_inline, _recover_claims,
_release_claim, _in_callers_atomic_block and _renewable().
Worker._run_attempt is also private. It returns a bool indicating
whether the outcome was recorded. A subclass override that returns
None is treated as "outcome not recorded". This keeps its row excluded
from claim recovery; it does not itself change the row's outcome.
Threads sharing a Worker can claim concurrently. No lock is held across
claim_one() or an override of it, including subclass statements after
the base claim returns. A database wait in one claim does not block
another claim through a Worker lock. A claim made through run(),
run_once() or the base claim_one() counts as in flight until its
returned row is registered. A short per-Worker lock protects this
bookkeeping and is never held across a database call.
On PostgreSQL and MySQL, a run_once() call inside its caller's atomic block
does not hold up lease renewal for the Worker's other tasks; see the SQLite
write-lock limit in
Production.
Worker.run_once() and testing.run_tasks() raise a claim's original
error at once. They make no immediate recovery attempt and no pending
recovery read before a later claim, whether inside or outside a caller's
transaction. A claim that committed stays RUNNING and is left to the
reaper, which requeues it with the attempt spent or marks it LOST on
the final attempt. The next call does not find that landed claim while
it remains RUNNING.
Neither run_once() nor run_tasks() closes the failed claim's
connection for recovery or tests or discards idle connections in Django's
PostgreSQL pool. The caller's connection and session, including advisory
locks, temporary tables and SET values, are left as the claim left them.
With psycopg_pool 3.3.0 or later, the run() loop discards idle pooled
connections without testing them after a failed pass in which the driver
reports a connection held or opened by that pass as lost. Any other failed
pass leaves the pool alone. With psycopg_pool 3.2.x, the loop still tests
idle connections after every failed pass.
Recovery runs only in Worker.run(), the loop used by ox_worker, on
PostgreSQL with psycopg 3, pooled or not, MySQL and SQLite. Pending recovery
is checked at the head of a poll pass unless another claim of the same
Worker is in flight. In that case it makes no recovery read; if a claim
begins before the look compares the claim generation, it releases nothing.
If a claim begins after that comparison, the look still releases the
orphans it read, but not that claim's row. Either way, recovery stays
pending, that pass claims as usual, and the next pass looks again.
Continuous claims on a shared Worker can keep recovery pending beyond
LOCK_TIMEOUT; the next look that passes the in-flight check expires
the window and releases nothing. The rows remain subject to the reaper.
A failed recovery read ends that pass with worker_poll_failed and
claim_recovery set to "pending"; no new claim is made and no error is
raised to a caller. PostgreSQL with psycopg2 makes no recovery attempt.
ox_worker does not share its Worker: its loop claims and checks recovery
on one thread.
The loop makes at most one recovery attempt when stopping. Only this
stop-time look has recovery-specific waiting limits. It waits up to five
seconds for claims in flight on the same Worker. If that budget runs out,
or a claim begins during the look, it gives up with
worker_claim_recovery_failed. A look given up because a claim began
after its claim-generation comparison may already have released rows,
each logged as worker_claim_released, before logging
worker_claim_recovery_failed.
On PostgreSQL with psycopg 3, pooled or not, the stop-time look uses a
private connection and a five-second budget covering that claim wait,
connection establishment and every reply. Host name resolution is outside
that bound and can exceed it. The private connection does not inherit a
session-level lock_timeout from the worker's existing connection and
can use its whole budget waiting on a locked table.
On MySQL, the stop-time look uses a private connection with five-second connect, read and write timeouts, raising any shorter configured timeout to five seconds. These limits are per operation, not an overall deadline, and host name resolution is outside them. mysqlclient may extend a read to 15 seconds. On SQLite, database waiting during the stop-time look is bounded by the busy timeout, not by an overall recovery deadline. The five-second claim wait applies on MySQL and SQLite too. A recovery error does not prevent shutdown, and these limits do not bound shutdown as a whole.
A Worker used in a process forked after the Worker was created takes a
new worker id in the child, retaining any -<slot> suffix. An at-fork
hook re-identifies the child. For forks that run no at-fork hook, such as
uWSGI's default, a pid check does so at the child's first claim_one(),
claim, recovery look or run(). Until that check, the child's copy can
still report the parent's id. The parent keeps its id.
The child starts with empty claim bookkeeping and fresh locks. It keeps the inherited heartbeat path and continues touching it under its new id; a heartbeat warning names the child's id. If several children share that path, a fresh heartbeat proves that one writer is alive, not that every child is alive.
This identity handling applies, for example, to a module-level Worker used
for run_once() under gunicorn --preload or uWSGI without lazy-apps,
or to a launcher that forks before the Worker claims anything. Handing an
already-claimed row to a child for execution is unsupported: the row
retains the parent's id, so the child with its new id cannot renew the
lease, and the reaper can requeue it while the child runs it.
This identity handling does not make arbitrary native forks safe or
establish that inherited database connections, threads or other resources
are safe to use. Separate worker identities prevent Worker.run()'s
recovery looks in one process from releasing the other's running task and
causing its body to run twice. ox_worker --processes starts fresh
children rather than forking them.
Versioning
django-ox follows Semantic Versioning:
- Breaking changes to any public surface above require a major version.
They are called out in the changelog under a
ChangedorRemovedheading, with the migration step. - Minor releases add, they do not break. Patch releases fix bugs or update documentation and package metadata; they add no features and change no behaviour beyond bug fixes.
A minor release may add a status value. A process still on the previous
minor release can't read a task in the new status. get_result() and
refresh() raise ValueError on it. The admin shows its status as - and
has no filter for it. Bulk discards skip it, and queue_stats(), the
django_ox_tasks gauge and ox_health don't count it. Upgrade every
process that shares a database before anything writes the new status. The
release notes name the value, say how it reads through django.tasks, and
give the upgrade and rollback steps.
Pin accordingly: django-ox~=1.8.0 accepts patch releases only;
django-ox~=1.3 accepts the current major line.
Deprecation policy
When a public surface is going to be removed or changed incompatibly, and a compatible path exists, it is deprecated before removal rather than dropped outright:
- The deprecation is documented in the changelog and, where it can be,
surfaced at runtime (a
DeprecationWarningor amanage.py checkmessage). - A deprecated surface is announced in a minor release and removed no earlier than the next major release.
Security fixes are exempt. A surface that cannot be kept without leaving a vulnerability open may change in a patch release. That is documented in the changelog, and in a security advisory where relevant.
Supported Python and Django
Each django-ox release is tested against the matrix below in CI, on SQLite and PostgreSQL 16 across the grid and MySQL 8 on the oldest and newest corners; these are the supported combinations.
| Django 5.2 LTS | Django 6.0 | Django 6.1 | |
|---|---|---|---|
| Python 3.12 | tested | tested | tested |
| Python 3.13 | tested | tested | tested |
| Python 3.14 | not supported by Django 5.2 | tested | tested |
Django 6.0 and later ship the Tasks framework in core. The Django 5.2 legs
install the django-tasks backport and run the whole suite against it, which
is what django-ox[backport] pulls in.
Support for per-task policy declarations differs from support for the backend itself:
| Declaration | Django 5.2 with django-tasks 0.12+ | Django 6.0 | Django 6.1 |
|---|---|---|---|
Bare @task with django-ox |
supported | supported | supported |
@task(max_attempts=..., backoff=..., timeout=...) with django-ox |
supported | not supported | supported |
The backport dependency already requires django-tasks 0.12 or later.
On Django 6.0, policy keyword arguments raise TypeError when the module
is imported, regardless of the configured backend. Bare @task still
builds a PolicyTask under OxBackend, with all three policy fields set
to None, so backend and queue defaults apply.
The policy fields are provisional. See the declaration reference for validation, precedence and typing guidance.
Django 6.1 changed which databases the system checks run against. A command
that runs the full checks and does not name a database now checks every alias
in DATABASES. Checking a SQLite or MySQL alias opens a connection and runs
one query. The cost grows with the number of aliases, and every such command
pays it. An alias that cannot be reached ends the command.
ox_prune, ox_health in database mode, and ox_import_beat_schedules
name the alias they work on and pass it to the checks. ox_worker passes
an empty list, so no alias is checked. Its poll loop can keep retrying a
database that refuses connections. Startup work before the loop, such as
loading database-backed schedules, can still access the database.
The system checks that need a database don't run for ox_worker.
A SQLite build without JSON support fails fields.E180.
manage.py check --database <alias> reports it, as do ox_prune,
ox_health in database mode, and ox_import_beat_schedules; each exits
non-zero. A worker on that alias runs tasks anyway, because SQLite stores
those columns as text, and it logs nothing. On MySQL, Django's column-type
checks emit warnings, so check still exits 0.
A database with no django-ox tables is a separate case.
check --database <alias> does not report that: it exits 0 and reports no
issues. migrate --check --database <alias> is what exits non-zero, and it
prints nothing at all. A worker that reaches the poll loop logs
worker_poll_failed on every pass. Run migrate --check for the alias
before you start a worker on it. A configuration error still stops a worker
at startup, because those checks don't need a database.
ox_health --heartbeat-file selects a separate branch before command
checks run. It runs neither system nor migration checks, validates no
database alias, constructs no task backend and makes no database calls.
It does not need --skip-checks. Project startup still runs before the
command: AppConfig.ready() and other startup code must be database-free
if the probe needs to survive a database outage.
For commands that run system checks, pass --skip-checks, or pass
--database to manage.py check, as appropriate. Django 6.0 is unaffected
by the check-scope change, and so is an alias on PostgreSQL. The
changelog has the mechanism, what a router does and does
not fix, and why --database alone is not enough on a command that does
not pass it on.
The support floor tracks Django's own: when a Python or Django version
reaches end of life upstream, a later django-ox minor release may drop it,
announced in the changelog. Databases: PostgreSQL, SQLite and MySQL 8 are
tested in CI. MariaDB 10.6+ uses the same claim path, since Django's own
floor guarantees SELECT ... FOR UPDATE SKIP LOCKED there, but it is not
part of the tested matrix.