Skip to content

API stability

This page states what counts as django-ox's public API, how versions change, and which Python and Django versions are supported. It is a promise about compatibility, so you can pin django-ox with confidence.

Public API

These are the supported surfaces. Changes to them are versioned and announced in the changelog. This page, not a module's __all__, is the statement of what is public. A module's __all__ may export names this page does not list; those names are not public.

  • The backend path django_ox.backend.OxBackend, referenced as a string in the TASKS setting, QUEUES beside OPTIONS, and every OPTIONS key it reads: MAX_ATTEMPTS, LOCK_TIMEOUT, BACKOFF_INITIAL, BACKOFF_MAX, TASK_TIMEOUT, TASK_TIMEOUTS, TASK_TIMEOUT_GRACE, SCHEDULES (with its documented per-schedule keys), and WORKER_CLASS, the dotted path of the Worker subclass ox_worker runs.
  • The management commands and their flags: ox_worker, ox_prune, ox_health. ox_worker's exit codes: 0 after a drain, 130 on a forced exit, 75 when the worker recycles itself after a stuck task thread. A worker that finishes under --batch or --max-tasks drains and exits 0, even when task attempts or individual schedule dispatches failed. Task outcomes stay on their task rows; schedule-scoped failures are reported in logs and leave no tick or task row. Exit 0 means the batch finished, not that every schedule enqueued. An abandoned dispatch pass prevents normal batch-empty completion until a later pass completes; reaching --max-tasks still ends the run. An invalid --max-tasks, or either flag with --processes above 1, exits 1 before any worker starts. Under --processes, the supervisor exits 0 when every worker drained or recycled, 1 when a slot hit the restart cap, and otherwise with the first other non-zero worker code. A worker killed by a signal reports 128 + the signal number, following the shell convention.
  • The stored-schedule write functions django_ox.stored.create_schedule, django_ox.stored.create_schedules, django_ox.stored.update_schedule and django_ox.stored.delete_schedule.
  • The heartbeat-file protocol: with ox_worker, one process writes PATH; above one process, the supervisor writes PATH.supervisor and slot i writes PATH.i. The modification time is the signal; file contents are not read or written. Every expected file must be regular and have an age between zero and the configured maximum, inclusive. A passing check means the expected controlling loops have advanced recently, not that tasks are progressing. A launcher that constructs one Worker with heartbeat_file and then forks it leaves the children touching the same inherited path under their own worker ids. A fresh shared path proves that one writer is alive, not that every child is alive. The documented file-mode JSON fields are public too.
  • The system check IDs, including the django_ox.E0xx and django_ox.W0xx identifiers, which you may list in SILENCED_SYSTEM_CHECKS. The IDs are stable; the messages are not.
  • ox_health's exit codes: 0 when every enabled check passes, 1 when a check fails. In file mode, passing means every expected heartbeat file is fresh. A value the command itself rejects exits 1 as well, such as --max-age 0, --max-heartbeat-age 0, --processes 0, a refused flag combination, or a --database alias that isn't in DATABASES. A value argparse rejects exits 2, such as --max-age nonsense, --max-heartbeat-age nan, --processes 1.5 or --format bogus, and so does an unknown flag.
  • The claim filter hooks Worker.claim_filter_q() and Worker.claim_filter_sql(), and where their result is applied: the fragment is conjoined to the conditions the candidate select filters on, ahead of its ordering and its limit. The statement around it is not promised.
  • The timeout helpers django_ox.deadline() and django_ox.remaining(), callable from inside a task. deadline() answers with a wall-clock time, and remaining() with seconds measured on the same monotonic clock the worker enforces the deadline with, so the two can differ by the size of a clock correction. remaining() is the one the watchdog agrees with.
  • The metrics module django_ox.stats: queue_stats, ready_count, oldest_ready_age, throughput, failure_rate, last_claim_age, waiting_counts, the QueueStats dataclass, and DEFAULT_WINDOW, the trailing window the rate functions default to.
  • The Prometheus surface: django_ox.metrics.render_prometheus, render_openmetrics and collector, the view django_ox.views.metrics with the using argument it takes from the URLconf, the django_ox.urls module with its metrics route name, and the metric names and label names listed on the Monitoring page. METRIC_NAMES is that list in code; CONTENT_TYPE_PROMETHEUS and CONTENT_TYPE_OPENMETRICS are the content types the view serves. A scraped name is a contract with every dashboard that reads it, so a rename is a breaking change. Help text is not part of the contract.
  • The actions module django_ox.actions: retry, discard and expire_lease, their accepted states, and their return values; retry_many and discard_many, the selections they accept and their (changed, skipped) return. RETRYABLE_STATUSES and DISCARDABLE_STATUSES are those accepted states in code; UPDATE_CHUNK_SIZE is exported for reading and its value may change. The admin page that calls them is a convenience over this module; its layout is not a contract, the two action names are.
  • The bulk module django_ox.bulk: enqueue_many(task, calls), its (args, kwargs) call shape, the input-order return and the all-or-nothing write. INSERT_CHUNK_SIZE is exported for reading; its value may change.
  • The exceptions django_ox.exceptions.TaskAbandoned, recorded against tasks whose worker stopped reporting with no attempts left (it records the lost lease, not a cause of failure), and django_ox.exceptions.TaskTimeout, a TimeoutError raised inside a sync task at its execution deadline and recorded against a timed-out attempt. The timeout comes from the task's own declaration, then its queue's TASK_TIMEOUTS entry, then TASK_TIMEOUT. Async tasks are cancelled and their timeout is recorded with TaskTimeout.
  • The structured-log contract: the event names and stable extra keys documented on the Monitoring page. This includes the policy events and keys, even though the declaration API is provisional. It also includes worker_claim_released, worker_claim_recovery_failed, worker_claim_recovery_expired and worker_claim_release_refused, and their documented keys. These are stable, not provisional. The recovery keys include claim_recovery, old_epoch, new_epoch, refunded_attempts, attempts and worker_ids. On worker_poll_failed, claim_recovery is "pending" while a recovery read is owed, or null otherwise; with psycopg2 it is always null. On worker_claim_recovery_failed and worker_claim_recovery_expired, it is always "expired".
  • The testing helpers django_ox.testing.ImmediateBackend, django_ox.testing.DummyBackend and django_ox.testing.run_tasks. The backends accept policy declarations but do not enforce retries, backoff or timeouts. run_tasks() drains due queued tasks through the configured worker class, including retries and backoff. It does not enforce timeouts. run_tasks() is public and provisional.
  • The database schema of OxTask and OxScheduleTick, evolved only through shipped migrations.
  • django_ox.__version__.

The producer-side API uses the Tasks framework's @task, .enqueue() and get_result(). Bare tasks follow that framework's contract. django-ox also accepts the provisional policy declarations below.

Provisional task policy

django_ox.tasks.PolicyTask and django_ox.tasks.BackoffCallback, also available as django_ox.PolicyTask and django_ox.BackoffCallback, are public, provisional APIs. This includes the max_attempts, backoff and timeout fields accepted through @task by OxBackend, and their representation on result.task.

This surface follows the names and two-argument callback in Django's open new-features proposals #142 and #144. It makes no promise of compatibility with whatever Django core eventually ships. timedelta callback returns are supported; timeout declarations use integer seconds. Retry scheduling updates run_after, rather than preserving its original value as proposal

142 suggests.

See Configuration for validation, framework support and precedence.

Not public

Everything else is an implementation detail and may change in any release without notice. That covers the django_ox.worker.Worker internals, the cron parser (django_ox.cron), the row-to-dataclass conversion (django_ox.results), the schedule loader (django_ox.schedules), the supervisor behind --processes (django_ox.supervisor) and the hidden --worker-index flag it starts each child with, and any name starting with an underscore. The exact SQL a claim emits and the model's non-schema helper methods are not part of the contract.

django_ox.heartbeat is also an implementation detail, not a public Python API. The documented heartbeat filenames, modification-time meaning, freshness rule and ox_health JSON fields are public contracts.

In django_ox.tasks, only PolicyTask and BackoffCallback are public, with the provisional status above. MAX_ATTEMPTS_LIMIT and validate_policy are implementation details. In django_ox.testing, only ImmediateBackend, DummyBackend and run_tasks are public. run_tasks is provisional. django_ox._run_tasks and Worker._task_body are private implementation details. Timeout implementation classes and methods, including TaskTimeouts.enabled and TaskTimeouts.for_attempt(), are not public.

Worker implementation details

Worker internals remain Not public. This includes _handed_off, _unsettled, _fence, _claims_in_flight, _claim_generation, _claiming, _claim, _claim_inline, _recover_claims, _release_claim, _in_callers_atomic_block and _renewable().

Worker._run_attempt is also private. It returns a bool indicating whether the outcome was recorded. A subclass override that returns None is treated as "outcome not recorded". This keeps its row excluded from claim recovery; it does not itself change the row's outcome.

Threads sharing a Worker can claim concurrently. No lock is held across claim_one() or an override of it, including subclass statements after the base claim returns. A database wait in one claim does not block another claim through a Worker lock. A claim made through run(), run_once() or the base claim_one() counts as in flight until its returned row is registered. A short per-Worker lock protects this bookkeeping and is never held across a database call. On PostgreSQL and MySQL, a run_once() call inside its caller's atomic block does not hold up lease renewal for the Worker's other tasks; see the SQLite write-lock limit in Production.

Worker.run_once() and testing.run_tasks() raise a claim's original error at once. They make no immediate recovery attempt and no pending recovery read before a later claim, whether inside or outside a caller's transaction. A claim that committed stays RUNNING and is left to the reaper, which requeues it with the attempt spent or marks it LOST on the final attempt. The next call does not find that landed claim while it remains RUNNING.

Neither run_once() nor run_tasks() closes the failed claim's connection for recovery or tests or discards idle connections in Django's PostgreSQL pool. The caller's connection and session, including advisory locks, temporary tables and SET values, are left as the claim left them. With psycopg_pool 3.3.0 or later, the run() loop discards idle pooled connections without testing them after a failed pass in which the driver reports a connection held or opened by that pass as lost. Any other failed pass leaves the pool alone. With psycopg_pool 3.2.x, the loop still tests idle connections after every failed pass.

Recovery runs only in Worker.run(), the loop used by ox_worker, on PostgreSQL with psycopg 3, pooled or not, MySQL and SQLite. Pending recovery is checked at the head of a poll pass unless another claim of the same Worker is in flight. In that case it makes no recovery read; if a claim begins before the look compares the claim generation, it releases nothing. If a claim begins after that comparison, the look still releases the orphans it read, but not that claim's row. Either way, recovery stays pending, that pass claims as usual, and the next pass looks again. Continuous claims on a shared Worker can keep recovery pending beyond LOCK_TIMEOUT; the next look that passes the in-flight check expires the window and releases nothing. The rows remain subject to the reaper.

A failed recovery read ends that pass with worker_poll_failed and claim_recovery set to "pending"; no new claim is made and no error is raised to a caller. PostgreSQL with psycopg2 makes no recovery attempt. ox_worker does not share its Worker: its loop claims and checks recovery on one thread.

The loop makes at most one recovery attempt when stopping. Only this stop-time look has recovery-specific waiting limits. It waits up to five seconds for claims in flight on the same Worker. If that budget runs out, or a claim begins during the look, it gives up with worker_claim_recovery_failed. A look given up because a claim began after its claim-generation comparison may already have released rows, each logged as worker_claim_released, before logging worker_claim_recovery_failed.

On PostgreSQL with psycopg 3, pooled or not, the stop-time look uses a private connection and a five-second budget covering that claim wait, connection establishment and every reply. Host name resolution is outside that bound and can exceed it. The private connection does not inherit a session-level lock_timeout from the worker's existing connection and can use its whole budget waiting on a locked table.

On MySQL, the stop-time look uses a private connection with five-second connect, read and write timeouts, raising any shorter configured timeout to five seconds. These limits are per operation, not an overall deadline, and host name resolution is outside them. mysqlclient may extend a read to 15 seconds. On SQLite, database waiting during the stop-time look is bounded by the busy timeout, not by an overall recovery deadline. The five-second claim wait applies on MySQL and SQLite too. A recovery error does not prevent shutdown, and these limits do not bound shutdown as a whole.

A Worker used in a process forked after the Worker was created takes a new worker id in the child, retaining any -<slot> suffix. An at-fork hook re-identifies the child. For forks that run no at-fork hook, such as uWSGI's default, a pid check does so at the child's first claim_one(), claim, recovery look or run(). Until that check, the child's copy can still report the parent's id. The parent keeps its id.

The child starts with empty claim bookkeeping and fresh locks. It keeps the inherited heartbeat path and continues touching it under its new id; a heartbeat warning names the child's id. If several children share that path, a fresh heartbeat proves that one writer is alive, not that every child is alive.

This identity handling applies, for example, to a module-level Worker used for run_once() under gunicorn --preload or uWSGI without lazy-apps, or to a launcher that forks before the Worker claims anything. Handing an already-claimed row to a child for execution is unsupported: the row retains the parent's id, so the child with its new id cannot renew the lease, and the reaper can requeue it while the child runs it.

This identity handling does not make arbitrary native forks safe or establish that inherited database connections, threads or other resources are safe to use. Separate worker identities prevent Worker.run()'s recovery looks in one process from releasing the other's running task and causing its body to run twice. ox_worker --processes starts fresh children rather than forking them.

Versioning

django-ox follows Semantic Versioning:

  • Breaking changes to any public surface above require a major version. They are called out in the changelog under a Changed or Removed heading, with the migration step.
  • Minor releases add, they do not break. Patch releases fix bugs or update documentation and package metadata; they add no features and change no behaviour beyond bug fixes.

A minor release may add a status value. A process still on the previous minor release can't read a task in the new status. get_result() and refresh() raise ValueError on it. The admin shows its status as - and has no filter for it. Bulk discards skip it, and queue_stats(), the django_ox_tasks gauge and ox_health don't count it. Upgrade every process that shares a database before anything writes the new status. The release notes name the value, say how it reads through django.tasks, and give the upgrade and rollback steps.

Pin accordingly: django-ox~=1.8.0 accepts patch releases only; django-ox~=1.3 accepts the current major line.

Deprecation policy

When a public surface is going to be removed or changed incompatibly, and a compatible path exists, it is deprecated before removal rather than dropped outright:

  • The deprecation is documented in the changelog and, where it can be, surfaced at runtime (a DeprecationWarning or a manage.py check message).
  • A deprecated surface is announced in a minor release and removed no earlier than the next major release.

Security fixes are exempt. A surface that cannot be kept without leaving a vulnerability open may change in a patch release. That is documented in the changelog, and in a security advisory where relevant.

Supported Python and Django

Each django-ox release is tested against the matrix below in CI, on SQLite and PostgreSQL 16 across the grid and MySQL 8 on the oldest and newest corners; these are the supported combinations.

Django 5.2 LTS Django 6.0 Django 6.1
Python 3.12 tested tested tested
Python 3.13 tested tested tested
Python 3.14 not supported by Django 5.2 tested tested

Django 6.0 and later ship the Tasks framework in core. The Django 5.2 legs install the django-tasks backport and run the whole suite against it, which is what django-ox[backport] pulls in.

Support for per-task policy declarations differs from support for the backend itself:

Declaration Django 5.2 with django-tasks 0.12+ Django 6.0 Django 6.1
Bare @task with django-ox supported supported supported
@task(max_attempts=..., backoff=..., timeout=...) with django-ox supported not supported supported

The backport dependency already requires django-tasks 0.12 or later. On Django 6.0, policy keyword arguments raise TypeError when the module is imported, regardless of the configured backend. Bare @task still builds a PolicyTask under OxBackend, with all three policy fields set to None, so backend and queue defaults apply.

The policy fields are provisional. See the declaration reference for validation, precedence and typing guidance.

Django 6.1 changed which databases the system checks run against. A command that runs the full checks and does not name a database now checks every alias in DATABASES. Checking a SQLite or MySQL alias opens a connection and runs one query. The cost grows with the number of aliases, and every such command pays it. An alias that cannot be reached ends the command.

ox_prune, ox_health in database mode, and ox_import_beat_schedules name the alias they work on and pass it to the checks. ox_worker passes an empty list, so no alias is checked. Its poll loop can keep retrying a database that refuses connections. Startup work before the loop, such as loading database-backed schedules, can still access the database.

The system checks that need a database don't run for ox_worker. A SQLite build without JSON support fails fields.E180. manage.py check --database <alias> reports it, as do ox_prune, ox_health in database mode, and ox_import_beat_schedules; each exits non-zero. A worker on that alias runs tasks anyway, because SQLite stores those columns as text, and it logs nothing. On MySQL, Django's column-type checks emit warnings, so check still exits 0.

A database with no django-ox tables is a separate case. check --database <alias> does not report that: it exits 0 and reports no issues. migrate --check --database <alias> is what exits non-zero, and it prints nothing at all. A worker that reaches the poll loop logs worker_poll_failed on every pass. Run migrate --check for the alias before you start a worker on it. A configuration error still stops a worker at startup, because those checks don't need a database.

ox_health --heartbeat-file selects a separate branch before command checks run. It runs neither system nor migration checks, validates no database alias, constructs no task backend and makes no database calls. It does not need --skip-checks. Project startup still runs before the command: AppConfig.ready() and other startup code must be database-free if the probe needs to survive a database outage.

For commands that run system checks, pass --skip-checks, or pass --database to manage.py check, as appropriate. Django 6.0 is unaffected by the check-scope change, and so is an alias on PostgreSQL. The changelog has the mechanism, what a router does and does not fix, and why --database alone is not enough on a command that does not pass it on.

The support floor tracks Django's own: when a Python or Django version reaches end of life upstream, a later django-ox minor release may drop it, announced in the changelog. Databases: PostgreSQL, SQLite and MySQL 8 are tested in CI. MariaDB 10.6+ uses the same claim path, since Django's own floor guarantees SELECT ... FOR UPDATE SKIP LOCKED there, but it is not part of the tested matrix.