Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
1.8.0 - 2026-10-04
Upgrading to 1.8.0: No database migration is required.
Upgrade workers so invalid stored schedules cannot stop other schedules or
queued tasks. On SQLite, look for interval schedules with every_seconds
above 62,135,596,800. Check stored start and end times against years 1 to
9999 in the database's time zone, and check SCHEDULES entries too.
Correct readable invalid rows in the admin change form or with
update_schedule, pause them, or delete them. Pausing does not repair a
row; enabling it again is refused until it is corrected.
For the unreadable PostgreSQL year-10000 row, the admin changelist returns
HTTP 500 and lists no rows, including healthy ones. The change form and
Disable action also return HTTP 500. update_schedule(row, enabled=False)
raises DataError. Delete this row with
delete_schedule(OxSchedule(pk=schedule_pk)) or a plain SQL DELETE.
If you applied importer output from 1.2.0 through 1.7.0, check schedules
restricting both day fields and crontabs from a zone other than
TIME_ZONE. Those versions passed crontab fields through as stored and
judged zones by two sample offsets. Also check beat rows with microsecond
intervals, which were listed as below one second. Regenerate output with
this version to see these rows by name.
Security
- Fixed worker termination from an uncaught overflow in interval tick calculation (CWE-248). Versions 1.2.0-1.7.0 are affected. An interval above 62,135,596,800 seconds, combined with certain phases, ends every worker reading it on its first dispatch pass and after each restart. Queued tasks do not run, and other schedules do not fire.
Stored rows can trigger this on SQLite through create_schedule,
update_schedule, admin add or change access, or direct table writes.
PostgreSQL, MySQL and MariaDB columns reject intervals this large.
A person who can edit SCHEDULES can trigger the settings path, which
fails before database access and is not limited to SQLite. The overflow
was reproduced on SQLite; the other database conclusions for this
interval value come from code inspection.
Every write path now limits every_seconds to 62,135,596,800.
SCHEDULES entries whose every exceeds the limit are refused by
manage.py check and at worker start with django_ox.E002.
On 1.7.0 with SQLite, disabling the triggering interval row through the
admin Disable action or update_schedule(row, enabled=False) is a
working workaround. A worker never builds a disabled row.
- Fixed worker startup and dispatch failures caused by stored schedule bounds outside the supported date range. On PostgreSQL, an out-of-range stored bound could prevent every worker from starting and stop every dispatch pass of a running worker, leaving other schedules unprocessed. Tasks already queued still ran.
This failure was reproduced on every release from 1.2.0 through 1.7.0.
The per-release runs used create_schedule with an end time late on
9999-12-31 in a zone west of UTC, stored in year 10000 UTC. Constructing any Worker() then raised
DataError. These runs used PostgreSQL, Django 6.0.8 and Python 3.12,
not each release's own support matrix.
create_schedule, update_schedule and create_schedules now reject
these bounds before writing them.
Writers now reject schedule bounds outside years 1 to 9999 in the
database's time zone. They also reject times with a time zone when
USE_TZ is off.
Fixed
- Skip and log stored rows that the worker cannot read or compare, rather than allowing one row to stop the worker. This includes values the database driver cannot convert, impossible dates, bounds outside years 1 to 9999, PostgreSQL infinity or BC dates, and MySQL zero dates. Stored intervals above the ceiling, or with a phase that could put a tick before year 1, are also skipped. Other schedules and queued tasks continue to run.
A tick that cannot be derived, or bounds that cannot be compared, are
reported per schedule as schedule_dispatch_error for any schedule
source. Database outages still stop the pass; they are not treated as
bad rows. An unreadable change marker makes the worker read the table
in full at each SCHEDULE_RECONCILE_INTERVAL (default 60 seconds).
Skipped rows are logged once, then at most once a minute per row per
worker while they remain skipped.
-
Allow readable stored schedules that fail validation to be paused with
update_schedule(row, enabled=False)and no other changes. This does not repair the row; enabling it again is refused until it is corrected. The admin's Disable action also works on these rows. In 1.7.0, it returned HTTP 500 for any row that failed validation. These operations do not work on the unreadable PostgreSQL year-10000 row. -
Validate stored schedule integers against the range of the database that stores schedules. If that database has a narrower integer range than the default database, out-of-range
every_seconds,phase_secondsandstarting_deadline_secondsnow produce field errors in the creation APIs,update_scheduleand the admin form. Previously, validation could pass and the database write could fail. Database integer-range validation is unchanged for schedules on the default database. -
Report field validation errors for non-numeric values in
every_seconds,phase_secondsorstarting_deadline_seconds,phase_seconds=None, mixed naive and aware values instart_timeandend_time, or empty values ("",[],()or{}) inevery_seconds,starting_deadline_secondsorend_time. These inputs previously raisedTypeErrorincreate_scheduleand at the validation step inupdate_schedule. Callers handling those errors should now handleValidationError. Inupdate_schedule, non-numericevery_secondsandphase_secondsstill raise aValidationErrorwith no field at an earlier step; that behavior is unchanged. -
Read crontabs with Celery's grammar, including wrap-around ranges, names, steps and lists. Write equivalent django-ox fields, using
*for a whole range. Write stepped ranges asa-b/n, using*/nonly for a set of three or more values over the whole field.1-59/2 1-23/2 1-31/2 * *fits the cron column. Long crontabs that 1.7.0 printed can print when the command finds a verified expression within the 128-character column limit. If the canonical form is too long, try a covering form within that limit. List the row if no verified expression fits.
When the weekday is restricted, write a day-of-month field covering
every date of its selected months as *. List rows that still
restrict both day fields instead of translating them: Celery requires
both to match, while a stored schedule runs when either matches.
-
Compute interval durations as beat does, with
timedelta(**{period: every}), including weeks and milliseconds. For example, 2,000,000 microseconds translates to 2 seconds. A fractionaleverycan be stored only on SQLite; PostgreSQL and MySQL use an integer column. On SQLite, 0.1 days translates to 8,640 seconds. Python'stimedeltarounds fractions to the microsecond before the whole-seconds check. List intervals whose resultingtimedeltahas fractional seconds, durations below one second, unknown periods or non-numeric values instead of translating them. -
Compare crontab zone files with the timezone data for
TIME_ZONE. Identical aliases qualify; matching current offsets alone do not. List incompatible or empty zones instead of translating them. Where the settings requireTIME_ZONEto have UTC-identical data, distinguish a zone this Python cannot load from missing zone files that prevent comparison. Neither failure establishes that the data differs from UTC's. -
Decode task arguments as beat decodes them, rather than applying different JSON decoding rules during import. List rows whose arguments cannot be translated instead of printing calls for them.
-
Name rows that stop beat's whole-table scheduling pass in a header line. With
USE_TZandDJANGO_CELERY_BEAT_TZ_AWAREboth off, rows with a start time or expiry stored with a UTC offset are listed with a reason explaining the offset-bound problem. In django-celery-beat 2.9.0, their due checks raiseTypeError; the failure can stop scheduling for the whole table, not just those rows. -
Quote names and values with bounded output in importer diagnostics, so an unusually large stored value does not produce an unbounded diagnostic.
-
Fix lease renewal on MySQL when callers share a Worker and one calls
run_once()inside its own atomic block on the worker's database. In 1.7.0, the renewal statement waited on that call's uncommitted, locked row for as long as the transaction stayed open, blocking renewal for every other task on the Worker. Other tasks could lose their leases, be requeued and run twice. Otherrun_once()calls could return late, and arun()loop asked to stop could take longer to return.
Renewal now excludes the uncommitted claim, which the caller's
transaction protects until commit publishes its outcome. Other rows
are renewed as before.
Affected deployments should upgrade. PostgreSQL, pooled or unpooled,
two plain run_once() calls outside atomic blocks, and two separate
Workers were unaffected by this defect.
SQLite's single write-lock limit is unchanged: other connections' writes still wait while the caller's transaction is open. On MySQL, callers using autocommit off without an atomic block must still use an unshared Worker. Tasks inside a caller's atomic block must not commit behind Django's back through raw transaction-control SQL or implicitly committing DDL, since that exposes a claim whose lease is not renewed.
- django-ox's PostgreSQL pool recovery sweep no longer runs statements on idle connections when using psycopg_pool 3.3.0 or later. In django-ox 1.5.0 through 1.7.0, the sweep tested every idle connection and waited without limit for each reply. A connection that stayed open without answering could hold the poll loop before it read a stop, or hold a task thread before it wrote its outcome.
Changed
-
The interval ceiling also rejects previously harmless schedules, such as an interval above the limit with phase 0. Such a schedule fired at most once, with its next tick after the year 3939. "Run once now" now reports an invalid stored row as not runnable instead of enqueueing it.
-
Limit phase and starting deadline values to 86,399,999,999,999 seconds in
phase_secondsandstarting_deadline_seconds. SQLite previously accepted larger values that were then skipped at every read. ASCHEDULESeveryorphaseoutside the range of Python'stimedeltanow producesdjango_ox.E002instead ofOverflowError. -
Make the beat importer print one all-or-nothing
created = create_schedules([...])call. TheSCHEDULABLE_TASKSfragment contains only entries used by printed rows. Apply the batch after checking the output. KeepSCHEDULE_SOURCEin the sameOPTIONS; without it, no stored schedule ever runs. Applying the output remains subject to validation, permissions, concurrent changes and database errors. -
List translations whose behavior differs from beat under "Translated, with a difference from beat". Every interval is listed because stored schedules count from a fixed instant and Celery counts from the last run. Clock-change notices come from the actual tick pattern, including spring-forward gaps and django-celery-beat 2.9.0's hour filter, rather than from clock transitions alone.
With USE_TZ on, stored schedules run matching repeated clock times
on both passes, while beat runs them on the first only. Notices give the
next relevant date; the scan covers ten years from import. Check
these differences before applying output.
-
Stop printing start times in importer output. Past starts are dropped. Rows with future starts are listed instead: beat runs once when the start arrives, while a stored schedule waits for its next tick. Arrange those starts separately before switching schedulers.
-
List rows whose timezone settings, schedule rules, arguments or destination limits prevent translation instead of printing calls for them. This includes values the destination database would refuse or alter: NUL characters, surrogates, excessively deep JSON, oversized integers, names equal under the name column's comparison, and names already taken.
Check each named reason and resolve omitted rows before switching. Run the command with beat's Python environment, timezone data and Django settings; the command cannot verify that they match.
-
With psycopg_pool 3.3.0 or later, recovery drains the pool only when the driver reports a connection as lost. Idle connections are closed without testing them, and the pool opens replacements. Connections already checked out are closed and replaced when returned. A failed pass or outcome-write attempt in which no connection was lost leaves the pool alone, including one that could not get a connection or failed with its connection still answering.
-
In repeated tests of the real
manage.py ox_worker, every connection was made silent and SIGTERM was sent one second later. With psycopg_pool 3.3.3 and the default 30-second pool timeout, the fixed worker exited 0 after 34.4 seconds. With a 2-second pool timeout, it exited 0 after 6.4 seconds. django-ox 1.7.0 was still running after 40 seconds and needed a second SIGTERM, exiting 130. The fixed worker on psycopg_pool 3.2.8 also remained running after 40 seconds and exited 130 after a second SIGTERM.
These are observations, not shutdown bounds. The tests ran on one shared machine through a local relay that acknowledged traffic at the TCP level, not a network blackhole.
This change adds no shutdown deadline. On the drain-capable path, the
sweep no longer runs statements on idle connections, but stopping can
still wait for the pool's checkout timeout and the stop-time recovery
look. Lowering the pool's timeout shortens that checkout wait.
Statements in the poll loop, renewal, outcome writes and task-thread
connection probes can still wait without limit.
- Draining replaces healthy idle connections as well as lost ones. It can open more connections and take longer to resume work than the old sweep. These are observations from pooled PostgreSQL tests on 2026-10-04, not bounds:
| Scenario | django-ox 1.7.0 | Now |
|---|---|---|
| One lost connection, pool of 10 | 1 connection opened | 9 connections opened |
Every session ended with pg_terminate_backend, four tasks in flight, pool of 10 |
9 connections opened | 31 connections opened |
| Pool too small, outcome writes time out on checkout | 0 connections opened | 0 connections opened; identical outcomes |
CONN_HEALTH_CHECKS enabled, loop holds no connection when every session ends, pool timeout 3 seconds |
Back at work in 7.0 seconds | Back at work in 12.0 seconds |
All four task outcomes were written once in the session-termination test on both versions. The additional connection opens do not increase the configured pool size or change the worker's connection budget.
-
psycopg_pool 3.2.x keeps the sweep used in django-ox 1.7.0 and is explicitly excluded from the drain-based stop-hang fix. It still tests idle connections and can wait indefinitely for a silent connection. Tests on psycopg_pool 3.2.8 opened 1, 9 and 0 connections in the first three scenarios above, matching 1.7.0 with identical outcomes.
-
psycopg_pool 3.3.3 or newer is recommended. The drain-based fix requires 3.3.0 or later. Separately, versions before 3.3.1 can lose pooled connections when task timeouts interrupt pool checkout, eventually leaving none available. That issue also affects django-ox 1.7.0. Version 3.3.3 also fixes pool maintenance threads terminating after 24 hours without work.
Added
-
Add the public, stable batch creation API
django_ox.stored.create_schedules(rows, *, user=None). Pass a list of mappings with the fields accepted bycreate_schedule. Creation is all-or-nothing: the function returns a list ofOxScheduleinstances, or[]for an empty batch. Names must be unique within the batch. OneValidationErrorreports every validation failure by row index, name, field and message. Permission checks follow validation, withPermissionDeniednaming every denied row. The batch shares one clock reading and tells workers once. A concurrent name conflict raisesIntegrityErrorand rolls back the whole batch. -
Annotate database errors from batch creation with the row concerned, both during the database check and at the write, so callers can identify the failing input to
create_schedules. -
Add the beat timezone option
--beat-timezone ZONEtoox_import_beat_schedules. Supply the timezone the Celery app ran beat in when the crontab table has no timezone column, or whenDJANGO_CELERY_BEAT_TZ_AWARE=False. When the option was given but was not needed, print a comment directly under the header explaining why. Print it only when at least one crontab row was read. -
Add a warning for unavailable pool draining:
connection_pool_cannot_drain, at WARNING level, for pooled PostgreSQL when the installed psycopg_pool lacksConnectionPool.drain(). It reports that restart recovery is preserved through the legacy sweep but the drain-based stop-hang fix is not in effect. It does not refuse startup or change the pool. It includeseventanddatabase; startup records also includeworker_id.
1.7.0 - 2026-09-28
Upgrading to 1.7.0: No database migration is required. Upgrade workers to recover their own unconfirmed claims. Recovery runs only in Worker.run(), the loop used by ox_worker, on PostgreSQL with psycopg 3, MySQL and SQLite. Worker.run_once(), testing.run_tasks() and PostgreSQL with psycopg2 retain 1.6.0 behavior on claim errors. A mixed fleet with 1.5.0 and 1.6.0 workers was tested; recovery applies only to upgraded workers' own claims.
If a claim commits but its reply is lost, the loop looks for eligible rows on later poll passes, or once at stop, and returns them to READY with the attempt refunded within LOCK_TIMEOUT and before lease expiry. Recovery is not guaranteed: worker death, prolonged outages, late commit visibility or continuous shared-Worker claims can leave rows for the reaper, consuming the attempt and becoming LOST on the final attempt. Stop-time recovery has backend-specific waiting limits, not a shutdown deadline. With Django's PostgreSQL pool, budget up to max_size + 3 server connections per worker process during that look, rather than max_size + 2.
A Worker inherited across a fork takes a new child id. Fork before the Worker claims anything; handing an already-claimed row to a child is unsupported. This does not make inherited connections or arbitrary native forks safe. An inherited shared heartbeat path proves only that one writer is alive. The recovery log events join the stable log contract.
Fixed
- Recover claims that committed but whose reply was lost, only in
Worker.run()(the loop used byox_worker), on PostgreSQL with psycopg 3, pooled or not, MySQL and SQLite. In 1.6.0, such a task stayedRUNNINGuntil the reaper requeued it with an attempt spent, or marked itLOSTon its final attempt even though its body never ran. Upgraded workers running this loop look for their own unconfirmed claims on later poll passes, or once at stop, and return eligible rows toREADYwith the attempt refunded, withinLOCK_TIMEOUTand before the row's lease expires.
A pending recovery look runs at the head of a poll pass unless another
claim of the same Worker is in flight. In that case it makes no read;
if a claim begins before the look compares the claim generation, it
releases nothing. If a claim begins after that comparison, the look
still releases the orphans it read, but not that claim's row. Either way,
recovery stays pending, that pass claims as usual, and the next pass
looks again. Claims on a shared Worker remain concurrent, as in 1.6.0;
recovery bookkeeping uses a short lock never held across a database
call. A claim made through run(), run_once() or the base claim_one()
counts as in flight until its returned row is
registered, including subclass work after the base claim returns.
ox_worker is unaffected by shared-Worker deferrals: its Worker is not
shared, and its loop claims and checks recovery on one thread.
A failed recovery read ends that pass with worker_poll_failed and
claim_recovery set to "pending"; no new claim is made and no error
is raised to a caller. Recovery is retried on later passes until it
succeeds or the window expires. Continuous claims on a shared Worker
can keep recovery pending beyond LOCK_TIMEOUT; the next look that
passes the in-flight check expires the window and releases nothing.
The rows remain subject to the reaper throughout.
PostgreSQL with psycopg2 retains 1.6.0's behaviour, with no recovery.
Worker.run_once() and testing.run_tasks() also retain 1.6.0's
behaviour: they raise the claim's original error at once, with no
immediate or pending recovery attempt. A claim that landed stays
RUNNING for the reaper, and the next call does not find it while it
remains RUNNING. It is requeued with the attempt spent or marked
LOST on the final attempt.
Recovery does not cover claims made by workers that have not been
upgraded, workers that die before recovery, or database outages that
outlast LOCK_TIMEOUT. If the claim's commit becomes visible only
after the recovery look, it may escape recovery and be reaped normally,
consuming the attempt and becoming LOST on the final attempt.
The loop makes at most one recovery attempt when stopping, without
scheduling a recovery retry. Only this stop-time look has
recovery-specific waiting limits. It waits up to five seconds for
claims in flight on the same Worker. If that budget runs out, or a
claim begins during the look, it gives up with
worker_claim_recovery_failed. A look given up because a claim began
after its claim-generation comparison may already have released rows,
each logged as worker_claim_released, before logging
worker_claim_recovery_failed.
On PostgreSQL with psycopg 3, pooled or not, the stop-time look uses a
private connection and a five-second budget covering that claim wait,
connection establishment and every reply. Host name resolution is
outside that bound and can exceed it. The private connection does not
inherit a session-level lock_timeout from the worker's existing
connection and can use its whole budget waiting on a locked table.
With Django's PostgreSQL pool, this connection is outside max_size;
budget up to max_size + 3 server connections during the stop-time
look, including the private renewal and timeout-watchdog connections.
On MySQL, the stop-time look uses a private connection with five-second connect, read and write timeouts; a shorter configured timeout is raised to five seconds. These are per-operation limits, not one deadline for the whole attempt. Host name resolution is outside these timeout bounds, and mysqlclient may extend a read to 15 seconds. On SQLite, database waiting during the stop-time look is bounded by the busy timeout, not by an overall recovery deadline. The five-second claim wait applies on MySQL and SQLite too. psycopg2 makes no stop-time recovery attempt.
A recovery error does not prevent shutdown. These limits do not bound
shutdown as a whole. The pool health checks that the run() loop makes
after a failed pass and MySQL's pre-existing wait for Django's
reconnection during the claim's atomic-block exit can delay shutdown
before the stop-time recovery look is reached.
Changed
- Give a Worker used in a process forked after its creation a new worker
id in the child, retaining any
-<slot>suffix, so thatWorker.run()'s recovery looks in one process cannot release the other's running task. An at-fork hook re-identifies the child. For forks that run no at-fork hook, such as uWSGI's default, a pid check re-identifies it at its firstclaim_one(), claim, recovery look orrun(). Until that check, the child's copy can still report the parent's id.
The child starts with empty claim bookkeeping and fresh locks; the parent keeps its id. As in 1.6.0, the child keeps touching the inherited heartbeat path. It does so under its new id, and its heartbeat warning names that id. A shared heartbeat path proves that one writer is alive, not that every child is alive.
This identity handling applies, for example, to a module-level Worker
used for run_once() under gunicorn --preload or uWSGI without
lazy-apps, or to a launcher that forks before the Worker claims
anything. Handing an already-claimed row to a child for execution is
unsupported: the row retains the parent's id, so the child with its
new id cannot renew the lease, and the reaper can requeue it while the
child runs it.
This identity handling does not make arbitrary native forks safe or
establish that inherited database connections, threads or other
resources are safe to use. ox_worker --processes starts fresh
children rather than forking them.
Version 1.6.0 shared the id after a fork without a recovery-release
risk, because it had no recovery look.
Added
- Stable structured-log events
worker_claim_released,worker_claim_recovery_failed,worker_claim_recovery_expiredandworker_claim_release_refused, with their documented extra keys.worker_poll_failednow includesclaim_recovery, set to"pending"while a recovery read is owed, ornullotherwise; with psycopg2 it is alwaysnull.worker_claim_recovery_failedis emitted only for a failed stop-time recovery look inWorker.run(), withclaim_recoveryalways set to"expired".
1.6.0 - 2026-09-27
Upgrading to 1.6.0: No database migration is required. Update processes that call run_once(), and use the updated ox_prune command. ox_worker behavior is unchanged from 1.5.0. These fixes do not require replacing workers.
This release preserves the caller's transaction when run_once() times out and avoids lease-renewal stalls inside a caller's atomic block. A timeout during a database statement can still make run_once() raise and force the caller's transaction to roll back. It adds the public, provisional django_ox.testing.run_tasks() helper for tests. The helper raises RuntimeError when the calling thread's connection to the worker's database is outside an atomic block and another thread's connection holds a TestCase transaction on that alias. The guard does not apply inside the caller's own atomic block, to another thread's plain atomic() block, or after gc.freeze(). Those cases can still return [] silently. ox_prune now rejects cutoffs on the first day of year one.
Added
- Public, provisional
django_ox.testing.run_tasks()helper for draining due queued tasks in tests. KeepOxBackendin test settings to exercise claiming, task outcomes, retries and backoff without a worker process.TestCaseuses savepoints and commit callback emulation. UseTransactionTestCasewhen testing worker autocommit behaviour. The helper does not enforce timeouts or dispatch schedules and reconcilers. Userun_tasks()to run queued tasks inside a test's transaction and a real worker for timeout tests.
Outside an atomic block on the worker's database, run_tasks() raises
RuntimeError before claiming anything if another thread holds a
TestCase transaction on the same database alias. The guard does not
detect this transaction after gc.freeze(). That case can still return
[] silently. For independent event loops, use TransactionTestCase or
django_db(transaction=True); Django's own async TestCase methods
remain supported.
Fixed
-
ox_prunenow rejects cutoffs on the entire first day of year one with the existing out-of-rangeCommandError, on every engine and withUSE_TZon or off. WithUSE_TZ=True, early cutoffs could previously raise anOverflowErrorduring Django's datetime conversion on SQLite and MySQL, after read-only queries had run, but before any statement carrying the cutoff ran or any row was deleted. This depended on the connection zone's UTC offset being negative in year one, not today. Rejecting the whole day makes the answer independent of engine and zone. First-day cutoffs that previously exited 0 now exit 1, including on PostgreSQL, withUSE_TZ=False, and for later cutoffs on SQLite and MySQL. The same rejection applies with--dry-runand--format json. -
Skip lease renewal for
Worker.run_once()inside a caller's atomic block on the worker's database. In 1.1.0 through 1.5.0, a task that outlivedrenew_intervalcould delay return on SQLite and MySQL by up torenew_interval + 5 safter the task ended.renew_intervalisLOCK_TIMEOUT / 3, so 100 s at the default 300 s. Outcome recording was unaffected. Calls outside an atomic block on the worker's database still renew their leases, including when autocommit is turned off. No caller changes are required. -
Worker.run_once()timeout handling inside a caller'stransaction.atomic(), including Django'sTestCase. From 0.3.0 through 1.5.0, an attempt ending inTaskTimeoutreset and closed the calling thread's database connections. This rolled back caller rows and the task row, then left the connection in autocommit. Later writes could commit and become visible to later tests, followed by a teardownTransactionManagementError.
Connections already inside atomic blocks now remain under the caller's
transaction control. When the connection remains usable, the attempt's
outcome is recorded inside that transaction and rolls back with it. A
timeout during a database statement can prevent outcome recording
entirely. run_once() then raises TransactionManagementError or a
driver error, and the caller's transaction must roll back. During ORM
writes inside the caller's transaction, run_once() raised
TransactionManagementError in every measured run on SQLite, PostgreSQL
and MySQL; during reads, a driver error occurred in all 30 PyMySQL runs,
in 1 to 7 of 30 PostgreSQL runs, and in no SQLite run.
Upgrade if tests or other synchronous code call run_once() inside
atomic blocks. Worker pool behaviour is unchanged.
No settings changes or migrations are required.
1.5.0 - 2026-09-25
Upgrading to 1.5.0: No database migration is required. Replace every worker to apply schedule isolation, dead-connection outcome recovery and the timeout-watchdog fix throughout the fleet. Schedule isolation adds no system check. When using Oxpull, deploy django-ox and Oxpull only as the tested exact-pinned release pair.
If you used ox_import_beat_schedules, read the Security entry
before generating or applying output. Upgrade and regenerate saved output.
If you applied output from an affected version, inspect the settings entries
and schedules as described there.
Deploy the policy-capable releases to every host before adding per-task declarations. A 1.4.0 host cannot import an unguarded policy declaration. With an import-compatible module, a 1.4.0 worker honours the stored attempt budget but ignores per-task backoff and timeout. See the rolling-upgrade guidance.
New backoff or timeout declarations affect queued rows on their next
attempt. New max_attempts declarations do not replace their stored
budgets. result.task preserves the declared task class and policy fields,
with routing reconstructed from the row; it does not report the row's stored
budget.
On Django 6.1, django-stubs 6.1.1 does not type the forwarded decorator
keywords. Suppressing call-overload makes the decorated task's static type
Any, losing task argument checking. Strict mypy also requires suppressing
untyped-decorator: # type: ignore[call-overload, untyped-decorator]. The
Django 5.2 backport needs no ignore; adding one can fail checks for unused
ignores.
Security
- Fixed a quoting bug in
ox_import_beat_schedulesoutput. Versions 1.2.0-1.4.0 did not safely quote some stored text in printed code. Text in the beat table could become Python that runs when the output is applied. Only projects that ran those versions of the importer, applied its output and had relevant beat records someone could edit are affected, for example through Django admin beat permissions. Installing django-ox or running the command without applying its output does not trigger the issue. Resulting Python could remain insettings.pyand run whenever settings load, with each loading process's privileges, or run with the privileges of the process applying generated schedule calls. Version 1.5.0 quotes every stored value withascii(); tests cover every location across SQLite, PostgreSQL and MySQL. Upgrade before generating output. Regenerate saved output from affected versions and review it before applying. If old output was applied, inspect pasted settings code, deployed and repository copies, retained schedule-call output and created schedules for unexpected Python, tasks, arguments or timing. If unexpected Python is found or its execution is suspected, investigate it as a security incident. These checks cannot rule out prior execution; upgrading or removing unexpected code does not undo it. Found during the project's own review; there are no reports of this issue being used.
Added
- The task admin has a Queue overview page linked from its change list. It compares retained rows by queue without adding aggregation queries to ordinary change-list visits. The page shows status counts, eligible backlog and age, five-minute throughput and failure rate, and time since the last claim.
ox_worker --batchexits once an error-free poll pass finds nothing to claim and began with none of its own tasks in flight, and--max-tasks Nexits after N claimed attempts, for cron and job runners. Both drain and exit 0, logworker_batch_emptyorworker_max_tasks_reachedwith theclaimedcount, and are rejected with--processesabove 1. Schedule-scoped failures do not hold a batch open. An abandoned dispatch pass prevents normal batch-empty completion until a later pass completes; reaching--max-tasksstill ends the run. Exit 0 means the batch finished, not that every schedule enqueued.Workertakes matchingbatchandmax_taskskeyword arguments, passed by the command only when the corresponding flag is given.schedule_dispatch_recoveredreports when a schedule that failed on this worker commits a tick again, by enqueueing a task or recording a first-sighting anchor. It carries the number of failed attempts infailures.ox_worker --heartbeat-file PATHoptionally updates local heartbeat files at the head of each poll and drain pass. A passing probe means the expected controlling loops have advanced recently, not that tasks are progressing. Single-process workers writePATH; with--processes Nabove one, the supervisor writesPATH.supervisorand children writePATH.0throughPATH.(N-1). The directory must already exist, be writable and be private to the container.ox_health --heartbeat-file PATHchecks every expected local heartbeat file instead of the database.--processes Nmust match the worker;--max-heartbeat-agedefaults to 60 seconds. File mode reads metadata only, runs no system or migration checks and makes no database calls. Project startup must also be database-free for the probe to survive a database outage. Database and queue options cannot be combined with file mode. Text and JSON output are supported; database-mode output, checks and--skip-checksbehavior are unchanged.heartbeat_write_failedandheartbeat_invalidate_failedwarning events report failed heartbeat updates and failed child-file removal. Update failures are nonfatal and retried every pass. Warnings are emitted once per writer/path or slot path, respectively.Workeraccepts the keyword-only argumentheartbeat_file, defaulting toNone. The command passes it toWORKER_CLASSonly when--heartbeat-fileis enabled. Existing fixed-signature constructors remain compatible without the flag; enabling it requires support for the keyword, otherwise the command raisesCommandError.- Provisional per-task retry and timeout policy through
@task(max_attempts=..., backoff=..., timeout=...)on Django 6.1 and Django 5.2 with django-tasks 0.12+. Django 6.0 remains supported, but its decorator does not accept these keywords.PolicyTask,BackoffCallbackand the three policy fields follow Django new-features proposals #142 and #144 without a compatibility promise toward the eventual Django core API. - Synchronous backoff callbacks can stop retries or choose an immediate
or delayed retry. Invalid results and callback failures log
task_policy_errorand fall back to the worker's exponential backoff. Callbacks are not time-bounded. -
Public policy-aware test backends:
django_ox.testing.ImmediateBackendanddjango_ox.testing.DummyBackend. They validate declarations but do not enforce retries, backoff or timeouts.task_policy_inertwarns on the first enqueue of each declaring task per backend instance. -
Two outcome-persistence events:
task_outcome_reconnected(WARNING) reports that, after a connection-level failure, a new connection recorded the outcome or confirmed that the first write had already committed. It carriesoutcome,already_writtenandduration_ms; the usual outcome log follows. Recovery retries outcome persistence at most once, never the task body.task_outcome_unrecorded(ERROR) reports that recovery on a new connection also failed. It carriesdropped_statusandduration_ms, with a traceback of the second failure. The outcome is unconfirmed, not necessarily absent.
Changed
schedule_dispatch_failednow reports an abandoned dispatch pass. Schedule-scoped failures, including database rejections, are reported asschedule_dispatch_errorat ERROR on every engine. Both events includedatabase,error,failuresandsuppressed. The first failure has a traceback; continued failures produce at most one summary per 60 seconds. Schedule reports are limited per worker, database alias and schedule; abandoned-pass reports are limited per worker. Alert on both events and readfailuresandsuppressedrather than counting log lines. For stored schedules, also alert onschedule_row_skippedandschedule_source_unavailable, regardless of batch exit.- Container probe recipes use optional local heartbeat files for
loop-liveness checks. Queue backlog, oldest age and last-claim age stay
in fleet alerting. Replace copied liveness recipes that use
ox_health --worker-timeoutor database-onlyox_health; an idle queue can fail the former, another worker's claims can mask a wedged worker, and a database outage can fail the latter across the fleet. - Probe guidance distinguishes loop liveness from task progress. Heartbeats cover the poll, drain and supervisor loops. Queue-age alerting monitors task progress, and configured task timeouts recover stuck slots. A hung database can make every heartbeat stale and cause a restart storm when failures trigger restarts. Examples include startup allowances and explicit thresholds to tune for the deployment. Plain Docker/Compose health checks report container health; restart behaviour requires separate supervision. The systemd unit retains process supervision.
- Attempt budgets are resolved at enqueue and remain stored on each row. Backoff and timeout use the worker's live task declaration on each attempt, including for rows already queued.
result.taskpreserves the declared task class and policy fields, withpriority,backend,queue_name,run_afterandtakes_contextreconstructed from the row. Re-enqueueing uses the task's declaredmax_attempts, or the backend's currentMAX_ATTEMPTSwhen none is declared.- Per-task timeouts take precedence over queue and worker defaults. Every attempt receives a fresh deadline; timeout grace and worker recycling remain worker-wide.
- Backend
MAX_ATTEMPTSvalues that cannot be converted, convert to a negative budget, or exceed the destination database's storage ceiling now produce system-check errordjango_ox.E011. Backend construction retains the configured value; enqueue and worker startup validate it when read and refuse invalid values, including startup with--skip-checks. Previously working values retain their behaviour under the deprecation policy below. OxBackend.max_attemptsis now read-only. Tests that assigned or patched this attribute must configureTASKSwith Django'soverride_settingsand obtain the backend under that override.- Worker startup validates fallback backoff values even with
--skip-checks. Explicit worker constructor overrides still accept zero-delay backoff; backendBACKOFF_INITIALandBACKOFF_MAXoptions must remain positive. - Operator retry refuses rows whose attempt count has reached 32767.
retry()returnsFalse; bulk retry counts those rows as skipped. - Structured retry logs include
retry_in_s. Terminal failure logs includereason, distinguishing exhausted attempts from a backoff callback that declined another retry. -
PostgreSQL pool-size warnings budget two unpooled connections per worker as a worst case, including a possible watchdog connection for a task-declared timeout. This is a capacity-planning assumption, not a count of open connections or a startup requirement.
-
Connection-level outcome-write failures whose recovery on a new connection also fails are now reported as
task_outcome_unrecordedinstead ofworker_error. Update alerts that rely onworker_errorfor these failures to also monitortask_outcome_unrecorded.
Deprecated
- Previously working backend
MAX_ATTEMPTScoercions and non-portable budgets now produce system-check warningdjango_ox.W004. Numeric strings, integral and fractional floats, and bools retain their oldint()conversion. Zero retains its one-attempt behaviour. Budgets above 32767 remain accepted where the database previously stored them: through 65535 on MySQL in strict mode and through 9223372036854775807 on SQLite. SetMAX_ATTEMPTSto a non-bool integer from 1 to 32767 before a future major release refuses these values. This is a system-check warning, not a runtimeDeprecationWarning. The new per-taskmax_attemptsfield is strict from introduction.
Fixed
ox_prune --older-thannow rejects durations too large to convert or subtract from the current time with aCommandErrornaming the value, before deleting any rows (#88).- The monitoring documentation lists the
worker_classstructured log key onclaim_filter_sql_missing. - On PostgreSQL and MySQL, dispatch continues to later schedules after a database rejection when rollback succeeds and the same database session remains usable. This includes non-finite floats in arguments and stored values that pass form validation but fail at enqueue. SQLite already continues to later schedules after such a rejection. Failed rollback, an unusable connection or a database error escaping a shared dispatch read still abandons the pass.
- On PostgreSQL, failed SQL in a
task_enqueuedreceiver is isolated to its schedule when rollback succeeds and the same connection remains usable. Integrity errors on the tick insert that are not genuine duplicate-key races are reported rather than silently ignored. - Stored schedules are traversed in primary-key order after settings
schedules in
SCHEDULESinsertion order. An edit does not change traversal order through PostgreSQL row movement. The order is for reproducibility; the tick unique constraint still coordinates workers. - Terminal failures from a task module that raises during import now
retain the
FAILEDoutcome and logtask_failed, including when the import exception is notImportError. An attempt that never rebuilt its task does not import it again or sendtask_finished. Attempts that rebuilt the task reuse it for the terminal signal. - The worker recovers outcome recording after a task encounters a dropped database connection, fixing a defect present in releases 0.1.0 through 1.4.0. If the connection loss is first detected during the outcome write, the worker retries recording once. Recovery is limited to outcome persistence and preserves lease fencing, a single error record and a single application of backoff. With Django's PostgreSQL pool, closing an unusable lost connection also sweeps that alias's open pool before the first outcome write or its single retry. The sweep checks idle connections and schedules replacements for dead ones. A failed poll pass also sweeps the pool before the usual poll-interval delay. Sweeps are best-effort and can block on connection checks; a subsequent checkout can still return a dead connection. Outcome recording still makes at most two writes, and failed poll passes are not replayed. An outage spanning both writes can leave the outcome unconfirmed; a row still requiring recovery is handled by the reaper after lease expiry. Execution remains at-least-once. Reconnection recovery runs with autocommit enabled and outside caller-owned transactions.
- For aliases without Django's PostgreSQL pool, the worker closes the
timeout watchdog's connection after every batch of stuck-attempt records,
including failed batches. This prevents stale-connection reuse after a
database restart or other disconnect between batches. Previously, a later
batch could fail to record
TaskTimeoutand backoff, leaving recovery to the reaper after lease expiry. The 1.4.0 release notes overstated cleanup by saying the connection was always closed or returned at the end of each batch. The pooled lifecycle is unchanged. An outage during a batch can still prevent recording. ox_import_beat_schedulesnow lists one-off, expired and empty-window rows under "Not translated, and why:" instead of importing them as recurring or live schedules. Supported schedules retain start times only when they are later than the import instant, otherwise starting when created; exclusive expiry bounds are always preserved by settingend_timeone microsecond earlier. If you applied output from versions 1.2.0-1.4.0, review imported schedules for unintended recurrence, missing start times and missing expiries; disable or correct affected schedules. Database read errors, non-null date bounds decoded asNone, and date-bound conversion failures stop the import with a concise error before any code is printed. Rows with invalid JSON arguments or non-finite numeric values are skipped with a reason. The application notes describe naive-local date interpretation and the difference in first-run behavior for tasks with a start time. Regenerating output with 1.5.0 does not resolve the cron day-field mismatch: Celery requires bothday_of_monthandday_of_weekto match, while django-ox accepts either, so review and manually adapt schedules that restrict both fields before applying them.
1.4.0 - 2026-09-23
If you use Django's PostgreSQL pool, check PostgreSQL max_connections
and role connection limits before upgrading. 1.4.0 workers open renewal
connections outside the pool, in addition to max_size. Task timeouts (TASK_TIMEOUT set, or any
TASK_TIMEOUTS value not None) also run the timeout watchdog, which opens
another connection outside the pool. Each worker process can hold up to
max_size + 1 server connections, or max_size + 2 with task timeouts.
No migration. For oxpull installations, use oxpull==1.4.0,
which pins django-ox==1.4.0.
PostgreSQL pool correction
This release fixes lease-renewal starvation affecting 0.2.0 through 1.3.1. The watchdog path is affected from 0.3.0.
The defect affects workers using DATABASES[alias]["OPTIONS"]["pool"]
whose effective max_size per process is below concurrency + 2, or
concurrency + 3 with task timeouts. "pool": True means max_size 4:
concurrency 3 or more is affected, or 2 or more with task timeouts.
Unpooled PostgreSQL, MySQL, and SQLite were not affected.
Leases can expire while task bodies still run. The reaper can reclaim that work and start another attempt. Bodies can run again, and tasks can end as FAILED or LOST. The defect was reproduced on 1.3.1. Earlier versions were identified by code inspection.
To investigate past impact on pooled PostgreSQL, look in the 1.3.1 logs
for the message text "Reclaimed stuck task" and "lost its lease"
(task_reclaimed and task_lease_lost in structured logs). Where
structured fields are available, check whether the reclaimed task's
held_by names a worker that was still running; worker_id on
task_reclaimed names the reaper. The message "Lease renewal failed"
(lease_renew_failed in structured logs) with a pool-timeout traceback
containing "couldn't get a connection after N sec" confirms the cause.
Plain-text handlers may omit structured fields, so zero hits for the
structured-log keys do not rule out past impact.
These events have no fixed order. With leases well above the pool's
30-second wait, such as the 300-second default, lease_renew_failed
comes first. With shorter leases, task_reclaimed comes first.
With very short leases, lease_renew_failed may not appear at all.
Affected tasks can end SUCCESSFUL after two or three body runs, or end
FAILED or LOST. queue_stats() returns counts per queue and status, not
rows; LOST counts alone miss most affected tasks and do not establish
the cause.
When opening a private connection fails, renewal and the watchdog try the pool with a checkout wait capped at 100 ms, then return borrowed connections. When no private renewal connection is open and lease time is short, renewal tries the pool first.
For 1.4.0, use at least concurrency + 1 pooled connections per worker
process for task threads and polling, plus the outside-pool server budget
above. Include every process, alias, other client, and reserved slot.
Pool fallback needs a spare pooled connection. It adds resilience, not
capacity. Add and budget a pooled spare if fallback must work under full load.
Give old workers pools of at least concurrency + 2, or concurrency + 3
with task timeouts, before rollout, or disable their Django pool.
Keep that budget until every old worker has stopped. This is also the
workaround if you cannot upgrade yet.
See Database connections and PostgreSQL pooling.
Added
ox_prune --format jsonprints prune counts as one JSON object.manage.py checkemitsdjango_ox.W003whenBACKOFF_INITIALandBACKOFF_MAXare both explicitly set to valid numbers and the initial delay exceeds the cap. Retries still run and waitBACKOFF_MAX.connection_pool_too_smallwarns at WARNING level when the worker alias's effective pool maximum is belowconcurrency + 1. It runs once perWorker.run(). It does not resize the pool or refuse startup.- Pooled PostgreSQL renewal reports these events:
lease_renew_degraded: WARNING on entering degraded renewal.fallback=succeededmeans renewal is using the pool; check server slots and connect stalls.fallback=failedmeans that tick did not renew leases. Alert on this event;lease_renew_missedcan arrive after live work has already been reclaimed.lease_renew_fallback: DEBUG for later successful pooled renewals.lease_renew_missed: WARNING with missed counts, at most every 30 seconds. After two consecutive misses, the next tick lands at the lease boundary; reclaim and another run are possible.lease_renew_recovered: INFO when private renewal resumes.watchdog_connection_unavailablewarns at WARNING level once per watchdog pass when neither connection path is available. A pass handles the stuck attempts recorded together. Their outcomes go unrecorded; recycling proceeds.
Changed
- Renewal ticks are scheduled start-to-start on every path, including unpooled
workers. An overrun starts the next tick immediately and resets the anchor.
Ticks do not overlap or catch up missed slots. Idle ticks still call
renew_leases()once, including custom overrides; the stock idle method opens no connection. - On pooled PostgreSQL, connection-acquisition failures use the new events
instead of
lease_renew_failed. Update alerts that relied on that event alone. Renewal statement failures still uselease_renew_failed. Unpooled failure reporting is unchanged. watchdog_erroralso reports failures while closing or returning a watchdog pass's connection, except that a database error while closing the private connection is suppressed without that event.- The schedule admin's Last tick column uses a subquery instead of a query per row. A 100-row page runs 5 statements instead of 105. The displayed value and database alias are unchanged.
Fixed
- With pooled PostgreSQL, lease renewal and the watchdog use private connections outside Django's pool, so pool exhaustion alone does not block them. If a private connection is unavailable, they fall back to the pool. Private connections need server capacity; pooled fallback still needs a spare pooled connection.
- The watchdog runs at most one acquisition sequence per pass, including attempts whose grace expires during acquisition or recording. It closes or returns the connection at the end without reconnecting between records. Stuck-attempt eligibility and recycling are unchanged.
- Unknown
ox_worker --backendaliases are rejected before workers start. The error names the invalid alias and configured choices. Ordinary invocation reports one error line;--tracebackretains the traceback.
Limits
Nothing needs setting for renewal or the watchdog. The private renewal
connect budget is the minimum of 5 seconds, the renewal interval, and a
positive OPTIONS["connect_timeout"]. The watchdog uses 5 seconds or a
smaller positive configured timeout. To bound task and poll-loop connects,
set OPTIONS["connect_timeout"] on the database alias.
One deadline covers all hosts. A stalled first host can leave later hosts untried. Synchronous DNS can exceed the deadline. Django's post-connect setup queries and renewal or recording statements are not bounded by it. These are not whole-tick or recycling deadlines.
The startup warning checks only the worker alias, not available server slots. Runtime degraded warnings report private-path failures that startup cannot detect.
Undersized pools can still cause task-query and outcome-write failures. After a PostgreSQL restart, tasks running at that moment can run again, including on 1.3.1; this release does not change that. Retries and reclaims can repeat side effects.
Task, polling, task-thread outcome-write, and hook connection paths are unchanged. PgBouncer and third-party pools were not tested.
Documentation
- The
db_workermigration notes explain which options do not carry over, how to select one queue with--queues default, and why omitting--queuesselects every configured queue. - Configuration docs clarify the renewal interval and per-queue timeouts.
1.3.1 - 2026-09-20
No code change since 1.3.0 except the version constant. No migration. Upgrade is optional.
Changed
- Update the README and package summary on PyPI.
Documentation
- Add the background tasks guide, "Why django-ox", and "Use cases".
- Update the benchmarks page with measurements against django-tasks-db 0.13.0 and recovery after worker death.
- Update the comparison page and correct the huey entry.
- Document how to order Oxpull Pro.
1.3.0 - 2026-09-17
One migration ships with this release. 0008_waiting adds a status
choice and runs no SQL. django-ox never puts a task into the new status by
itself. An install that doesn't use workflows in Oxpull Pro upgrades and
rolls back as it always has.
A 1.2 process doesn't know the new status. Its admin shows a waiting task's
status as - and has no Waiting filter. Its discard, discard_many and
Discard selected tasks action skip waiting rows. Its queue_stats(),
django_ox_tasks gauge and ox_health don't count them. Its get_result()
and refresh() raise ValueError on a waiting task. Nothing in 1.2 releases
one.
Rolling back to 1.2 after workflows have run. A WAITING task is a task
row whose status is WAITING. No worker claims it, the reaper doesn't reap
it, ox_prune doesn't delete it, and retry skips it. django-ox writes that
status onto no row of its own, and releases no row that holds it. Something
built on django-ox does both, and on this release that is workflows in Oxpull
Pro.
Turning workflows on, and turning them off again, is Oxpull Pro configuration, and Oxpull's own documentation has those steps. This section is the django-ox half. See Pro.
migrate django_ox 0007 refuses while any task is WAITING on the database it
migrates. A version from before 0008_waiting can't read a waiting task and
has nothing that would release one, so those rows would sit there for good.
The refusal reads through the connection being migrated, so waiting rows on
one alias neither block nor excuse the way back on another.
Take these steps in order, for each database that holds django-ox's tables.
- Stop whatever writes WAITING tasks from writing more. Then wait for the requests, jobs and transactions that were already writing one to end.
- Finish or cancel the work the remaining waiting rows belong to. Only what wrote a waiting row releases it, so django-ox can't do this part for you.
- Stop every process that can write a WAITING task, and keep it stopped until it runs 1.2. That's any process that enqueues, such as web and ASGI processes, enqueue-only services, workers and cron jobs. A process that's already running keeps the settings it started with, so a settings change doesn't stop it.
- On every database alias, run
OxTask.objects.using(alias).filter(status="WAITING").count(). Each count must be 0. If one isn't, don't migrate. With the processes from step 3 stopped, the rows left belong to work that hasn't finished or been cancelled. Finish or cancel it and count again. A count that goes up between two runs means a process that writes WAITING tasks is still running. Find it and stop it first. - Run
migrate django_ox 0007 --database aliasfor each alias. It refuses while any task on that alias is WAITING. - Deploy 1.2 everywhere. Then start the processes you stopped in step 3.
The count and the migration look at the rows that exist when they run. Neither
stops a process on this release from writing a WAITING task afterwards, and
nothing at 0007 refuses one. That's why the processes from step 3 stay
stopped until they run 1.2.
A backup that holds a WAITING task is not refused by a 1.2 install.
Nothing checks, and the row restores. A 1.2 install then leaves it where it
is. No worker claims it, ox_prune does not delete it, and retry and
discard both return False. get_result() raises ValueError, and
ox_health counts the row in no column and still reports OK. Put this
release back and the row moves again, so restore such a backup into this
release or a later one.
Django 6.1 runs the system checks against every database alias. A
command that runs the full checks and does not name a database now checks
every alias in DATABASES. Checking a SQLite or MySQL alias opens a
connection to it. An alias the machine cannot reach ends the command before
it does any work, and the usual case is a replica. runserver is one of
these commands, so a developer whose replica is unreachable cannot start the
dev server. A reachable alias is opened too, so each extra alias costs a
connection on every such command. Django 6.0 does not do this, and nothing
in django-ox changed.
django-ox's own commands don't check an alias you didn't ask for.
ox_prune, ox_health and ox_import_beat_schedules each pass the alias
they work on to the checks. ox_worker passes an empty list, so the checks
that take a database run against nothing: a worker has to start while its
database is down and wait for it. Both are new in this release, and both hold
whether or not you pass --database.
For every other command, --skip-checks is the cheapest way out and needs
no settings change. manage.py check has no --skip-checks. Give it
--database and name the alias you want checked.
--database on its own is not the exemption. A command has to pass the
flag to the checks, and most do not. showmigrations, sqlmigrate,
dumpdata and flush all accept --database and still check every alias.
SILENCED_SYSTEM_CHECKS does not help either. The connection raises before
there is a check message to silence.
A database router is what fixes the commands that name no alias, runserver
among them. How far it fixes them depends on the engine. Django skips an
alias whose allow_migrate returns false for the model. A router that keeps
django-ox's tables on one alias therefore stops those checks reading the
others for django-ox's models. Other apps' models are still checked on every
alias. On SQLite nothing but a JSONField column opens a connection, so the
router is enough where django-ox holds the only ones. On MySQL every field
check reaches the server, so contenttypes ends the command whatever the
router says. There the router has to send every app to one alias.
A router does not cover Django's backend checks for an alias you name.
manage.py check --database <alias> runs them for that alias, and no router
is read on that path. On MySQL those checks ask the server for its
sql_mode, which opens the connection. An unreachable alias named that way
ends the command, router or no router. It's the same read behind the
mysql.W002 warning under Changed. On PostgreSQL those checks open nothing,
so naming an unreachable alias costs nothing there. On SQLite they open
nothing either. The JSONField check below does, so an unreachable SQLite
alias still ends the command. It passes only behind a router that keeps
django-ox's tables off that alias. Pass --database only for aliases the
machine can reach.
Two field checks reach a connection. On SQLite, Django asks the alias
whether it supports JSONField, and the answer comes from a query. On
MySQL, the backend validates each field's column type, and reading that type
asks the server for its version. The first needs a JSONField column. The
second fires on the first field of the first model, whatever its type.
PostgreSQL ships no backend field validation. It answers the JSONField
question from a constant, so a PostgreSQL alias is unaffected.
Added
OxTask.Status.WAITING,stats.waiting_counts(), and thewaitingvalue of thestatuslabel ondjango_ox_tasks. A waiting task reads asREADYthroughdjango.tasks, which means it hasn't finished, not that a worker can take it. Workers never claim it,ox_prunenever deletes it, and retry skips it. It isn't backlog, soready_count(),oldest_ready_age()andox_healthleave it out.QueueStatskeeps its fields, and none of them counts a waiting task.ox_prune --queuerestricts pruning to one queue's task rows, matchingox_health --queue, so queues with different retention needs can each be pruned with their own--older-than. Old schedule ticks are still pruned for every schedule.ox_health --format jsonprints the check figures as one JSON object for container healthchecks and monitoring agents. On a failing check the object is still printed, and the exit status is unchanged.--databaseonox_prune,ox_healthandox_worker, naming the alias to work on. It defaults to the aliasOxTaskwrites to, the waymigrate --databasedefaults to one.ox_workernames it in the command line of each--processeschild, so one router answering differently in two processes cannot split a fleet across two databases.ox_prune,ox_healthandox_import_beat_schedulespass their alias to the system checks, which is whatmigratedoes;ox_import_beat_schedulesalready had the flag and now does this too.ox_workernames no alias there, for the reason under Changed. The flag is not checked against the router. A worker pointed at another alias works on that one, while the admin,statsand the actions still read the router's. Nothing warns.
Changed
retry_manyanddiscard_manysort the ids and lock each thousand rows in primary key order before they update them. That's one more statement per thousand rows on PostgreSQL and MySQL. A bulk retry or discard of rowsox_pruneis deleting now waits for it. Before, it could fail with a deadlock. The call locks every row it was given, whatever its status, until it ends.- When
retry_manyordiscard_manyopens its own transaction, a deadlock or a serialization failure starts the call again. It stops after three attempts in all. Inside a transaction of your own, the error still reaches you. ox_prunestill commits batch by batch. It now locks each batch's rows in primary key order. A batch that hits a deadlock or a serialization failure runs again in a new transaction, three attempts in all. If it still fails, the command exits non-zero. The batches before it stay deleted, and runningox_pruneagain deletes the rest.discardanddiscard_manyaccept WAITING, andDISCARDABLE_STATUSESincludes it.django_ox_taskshas awaitingsample for every queue, so a sum over itsstatuslabel now counts waiting tasks too.django_ox.stats,django_ox.metrics.collect(),render_prometheus(),render_openmetrics()andcollector()takeusingto name the alias to read. Left out, they read the aliasOxTaskwrites to. On a project with no database router that is the same connection they always used. The shipped endpoint takes it from the URLconf:path("ox/metrics", metrics, {"using": "replica"})serves scrapes from a replica and keeps them off the primary. It's a mount argument rather than a query parameter, so whoever scrapes can't choose the database.ox_workerhands the system checks no database alias, so starting a worker opens no connection before its first poll. A worker started while its database is down logsworker_poll_failedand polls again a second later, on Django 5.2, 6.0 and 6.1 alike, instead of exiting. That's the path it already takes when the database goes away while it runs, and a process manager restarting a worker into a database that's still down gives up long before the database is back. What that costs: the system checks that need a database don't run forox_worker. A SQLite build without JSON support failsfields.E180.manage.py check --database <alias>reports it, and so does every other django-ox command; each exits non-zero. A worker on that alias starts and runs tasks anyway, because SQLite stores those columns as text, and it logs nothing. The other case is a database with no django-ox tables.check --database <alias>does not report that: it exits 0 and reports no issues.migrate --check --database <alias>is what exits non-zero, and it prints nothing at all. The worker logsworker_poll_failedon every pass. A configuration error still stops a worker at startup, because those checks don't need a database.- On MySQL,
ox_prune,ox_healthandox_import_beat_schedulesrun Django's database checks for their alias on every invocation. On a connection without strict mode that printsmysql.W002each time, including from cron. Turn strict mode on, which is what the warning asks for, or putmysql.W002inSILENCED_SYSTEM_CHECKS. ox_prune,ox_healthandox_import_beat_schedulesreport a database they can't reach as one line,Database unreachable: <reason>, and exit non-zero.ox_health --format jsonprints its object with the figures null and the reason inproblems, wherever in the run the database was found to be down.ox_health --max-ageand--worker-timeoutaccept the duration formsox_prune --older-thantakes (7d,24h,90m,45s). A plain number still means seconds, fractions included.- A worker ended by a second stop signal exits with code 130 without
logging
second signal received; forcing exit.first; the--processessupervisor still logs its own line. - The source distribution no longer carries the repository's
.gitignore. It carries the package, the licence, the README, the changelog and the files that build it.
Fixed
ox_pruneandox_healthno longer read a replica. Under a router that sends reads to one, every query they made went there while their writes went to the primary.ox_prunedeletes in batches, and the loop ends when its candidate read comes back empty. A replica that is up and behind never comes back empty. So the command emptied the primary and kept going: 20 rows to 0, still running a minute later, and the operator had to kill it.ox_healthanswered from the replica. Over 40 READY tasks six hours old on the primary it printedOK: backlog=0and exited 0. With--format jsonit said"ok": true. A container healthcheck on it reported green over a queue that was stuck.ox_prunenow finishes and reports what it deleted;ox_healthreports the backlog the workers see. Present in every release from 0.1.0 to 1.2.0.
Both read the alias the task rows are written to, and so does the rest
of django-ox's reading of its own rows: django_ox.stats, the metrics
renderings and the endpoint, django_ox.actions, get_result(),
enqueue() and enqueue_many(), the worker's claim and its completion,
the reaper, the stored schedules, and both admins.
That list is what a test in the suite drives, under a router that refuses any such read on a replica. It drives the admin over HTTP, because calling an admin method is not the same as opening the page. For both models it opens the changelist and its filters, search, facet counts, show-all view and raw-id popup. It opens the history page and runs every action the admin offers. For schedules it also opens the change form's GET and POST, the add page and the redirect after it, and both delete confirmations. It calls the autocomplete endpoint another app's form uses to fill in a task. Those pages are what the test covers, and over them the sweep is a property the suite holds rather than a claim.
The admin reads the primary, and no setting changes that. Every
page of both admins reads the database its writes go to. Before this
release those pages followed the read alias, so a router that splits
reads sent them to a replica. On a replica that lags, the schedule
admin's add page then raised OxSchedule.DoesNotExist on the row it
had just written. The load now lands on the database your workers use.
If you were serving admin reads off a replica on purpose, this ends it.
That is the right default, because the admin is not a reading page. It writes back what it read. A change form submits every field, including the ones nobody touched, so a form built from a replica overwrites newer values on the primary. Nothing raises and nothing is logged.
What you can still point at a replica is what you ask for by name.
django_ox.stats, collect(), the renderings and the endpoint take
using, so path("ox/metrics", metrics, {"using": "replica"}) keeps
scrapes off the primary. Your own queries are untouched: a router you
wrote still sends your reads of the task table where you send them.
Configuration
covers what django-ox pins and two places it cannot reach.
- A worker no longer loses a task's result under a router that sends reads
to a replica. On MySQL and MariaDB every claim re-reads the row it has
just claimed, and that read followed db_for_read. The worker was handed
the row as it stood before the claim, so it ran the task holding a lease
the row no longer had: its own finish write matched no row, the result
was lost, and the row sat RUNNING until the reaper requeued it. The
re-read now names the alias the claim was written to. PostgreSQL claims
in one statement and reached this only through a subclass that overrides
claim_filter_q() without claim_filter_sql(). SQLite takes the
compare-and-set path and was never affected.
- Stored schedules read the database they are written to. Under a router
that sends reads to a replica, creating or editing a schedule checked the
name against the replica, so a name already taken on the primary passed
validation and the INSERT raised IntegrityError, which the admin showed
as a server error rather than as "Schedule with this Name already
exists." update_schedule() and the admin's add page then read the row
back from the replica: a row the replica had not seen yet raised
OxSchedule.DoesNotExist after a write that had succeeded, and an older
one came back holding the values that write had just replaced. The
Enable, Disable and Run selected schedules once now actions
read their rows from the replica too, so a manual run could carry
arguments that had already been changed. All of it now reads the alias
the schedules are written to. The unique index on the name is unchanged:
a check can't win a race against an insert that commits between the check
and the write, so the index is what makes the name unique and the check
is what turns the ordinary duplicate into a field error.
- The schedule admin no longer writes a stale copy of a row over a newer
one. Under a router that sends reads to a replica, Django built the
changelist and the change form from the replica. A change form submits
every field, so Save wrote the replica's values back over the
primary's. There was no error and no warning, and the newer values were
gone. update_schedule() takes the row lock and writes only the
submitted fields to stop this, and a stale form walked through it.
Two more followed from the same read. Save and continue editing on the add page looked the new row up on the replica. The person was told the schedule they had just made doesn't exist, and landed on the admin index. Delete selected schedules counted the rows on the replica, found none, and returned before it reached django-ox's own delete. It deleted nothing and said nothing at all.
Both admins now read the alias their rows are written to, and so does the
Last tick column. The task list, its filters and a task's page read
the replica too. A task enqueued a moment ago was missing from the list,
and a task's own page reported it as deleted. The schedule pages are
present since 1.2.0, which added stored schedules; the task pages since
0.3.0, which added them.
- ox_prune --include-failed no longer deletes a row that an operator retries
while its batch is being deleted. The DELETE matches rows by primary key
alone, so a FAILED or LOST row retried just before it ran was deleted anyway.
The retry still reported success, in the admin and in django_ox.actions. A
row discarded at that point was deleted too. Each batch is now checked again
inside the transaction that deletes it. Only rows that still qualify are
deleted, and nothing else can write to them until that transaction ends. A
retry that reports success now keeps its row. Without the flag, ox_prune
was not affected.
- retry_many and discard_many open their transaction on the database
that OxTask writes to. Under a router that sends OxTask to another
database, each UPDATE committed by itself. An error part-way could leave
some rows moved.
- The lease documentation put the renewal margin at two consecutive missed
renewals. It is one. The renewal loop waits LOCK_TIMEOUT / 3 after each
renewal rather than firing on a fixed schedule, so every round costs the
wait plus the UPDATE that renews. Three rounds therefore always come to
more than the lease, and a second miss in a row leaves the lease expired
and the task reclaimable. Nothing in the worker changes. Size
LOCK_TIMEOUT for one missed renewal.
- ox_health --max-age and --worker-timeout accepted nan, inf and
numbers that round to inf. A threshold set to one of those could never
be exceeded, so that check could not fail. It passed however old the
oldest task waiting to run or the last claim was. --max-backlog takes
an int and was not affected, and neither was the --worker-timeout
branch that reports no claim at all. They are refused as usage errors
now.
- A stop signal could leave an idle ox_worker hung instead of draining.
It stayed hung until a second signal or the process manager ended it,
or for good when it was a worker process whose supervisor had been
killed. A worker that has finished starting now drains on the signal.
- A worker process whose supervisor died while the worker was still
starting ran on as an orphan. It now drains and exits having claimed
nothing.
- If the --processes supervisor hit an error while running, a second
stop signal did not send SIGKILL to a worker process that would not
exit, and the supervisor waited for it forever. The second and third
signals now escalate as they do in any other stop.
1.2.0 - 2026-09-12
One migration ships with this release. 0007_oxschedule creates two
tables, django_ox_oxschedule and django_ox_oxschedulechange, with a unique
index on the schedule name and a check constraint. It touches neither the task
table nor the tick log, so there is no index build on a table workers are
reading: on PostgreSQL the locks are on the new tables only, inside one
transaction, and on MySQL each CREATE TABLE takes a metadata lock on its own
new name. Enqueues, claims and the dispatch of settings schedules carry on
through it, and a 1.1.0 worker keeps running beside a worker on this release
once it is applied. Migrate before starting any worker whose backend names
django_ox.stored.DatabaseScheduleSource, and before opening the schedule
admin: both read the new tables.
One thing a rolling deploy does not close. A settings schedule that has no tick yet and is first seen during the rollout can be anchored twice, once by a 1.1.0 worker and once by a worker on this release, when their passes hold different ticks: the 1.1.0 worker does not take the first-sighting latch this release adds, so the latch serialises only the new workers. The later of the two ticks is then recorded with no task and does not fire. That is the race 1.1.0 has between two of its own workers, not one this release introduces, and it ends when the last 1.1.0 worker stops. Once a schedule has any tick, both versions coordinate on the tick log's unique index and nothing anchors again. Stored schedules are not affected: a row carries its boundary and never anchors.
Migrating back to 0006 drops both tables and every stored schedule with
them. The tick log keeps its rows, including those named db:<id> for
schedules that no longer exist. A worker on this release that is still
running warns on every pass that it cannot read the schedule tables, and
keeps dispatching the settings schedules from the set it last read; stop it,
or migrate forward again. sqlmigrate django_ox 0007 prints the statements
for your engine.
Added
OPTIONS["SCHEDULE_SOURCE"], a dotted path to the class a worker asks for its active schedules on every dispatch pass. It defaults to readingOPTIONS["SCHEDULES"], so settings-declared schedules are unchanged. A source owns its own freshness, which is what lets one read somewhere that changes without the worker knowing.django_ox.E006reports a source that cannot be built or has noschedules()method.- Fixed-interval schedules. A
SCHEDULESentry takeseveryinstead ofcron, as atimedeltaor a number of seconds, with an optionalphaseto offset the sequence. Exactly one ofcronandeveryis required. Ticks are counted from a fixed instant rather than from the last run, so a restart, a pause or an edit cannot shift the cadence: every worker derives the same instants from the definition alone, which is what keeps dispatch leaderless.everymust be at least one second, because the dispatch loop cannot honour anything faster. - A tick whose instant has not arrived is not enqueued. On the day a zone springs forward, an hour of wall-clock labels never happens, and a label inside it resolves to an instant on the far side of the gap. That tick is held until its instant arrives and fires once.
- A registry of the tasks a schedule may name.
@schedulable("reports.daily")above@taskexposes a task under a key, andOPTIONS["SCHEDULABLE_TASKS"]does the same from settings, which is the only channel a system check can see. It is the boundary the stored schedules rest on: a row names a key the code owns rather than a dotted path anyone with the change permission could choose. Registration is discovered lazily on first use, so a project not using the registry imports nothing it did not already import. An entry may declare apermission, checked in addition to the model permissions before a schedule naming that key is written, when auseris passed; it is enforced indjango_ox.stored, not only in the admin.django_ox.E007reports a bad entry. django_ox.registry.ArgsForm, a Django form for a schedule's arguments that closes two defaults written for HTML posts rather than stored rows: an unknown argument is an error instead of being ignored, and a text field refuses a non-string instead of coercing it, so5cannot reach a task as"5".OxSchedule, a recurring schedule stored as a database row so it can be created, retimed and paused without a deploy, andOxScheduleChange, the single row a worker reads to know whether the stored schedules moved. A row names a registry key rather than an import path. It carriesstart_timeas its activation boundary, written when the row is created rather than when a worker first notices it, and records the timing and pause state that boundary was set for, so retiming a schedule reschedules it from the moment of the change instead of firing a tick that already passed. Because that record is derived from the row rather than incremented by a write path, a retime or a pause made with a bulkupdate()is noticed too, at the next read, and the boundary moves to that read: the worker remembers the boundary it found the row stale against, so a row re-enabled the same way before the move is written still gets it. It keeps that memory until the move commits, so a dispatch pass run inside a transaction the caller rolls back has forgotten nothing by its next one. However many workers find the row stale together, only the first of them moves the boundary: the row counts its boundary writes and each worker records the count it saw, because a heal writes the boundary to its own clock, and writing it is not the same as changing it. A change made and reverted between two reads is not, and the schedules page says what that means for a bulk pause and resume. Ticks are recorded against the row, asdb:<id>, so renaming a schedule keeps its history; a settings-declared schedule may not use that prefix, andmanage.py checkrefuses one that does.start_timeandend_timebound which ticks of a stored schedule fire, andstarting_deadline_secondsdrops a tick that is later than the deadline rather than running it however stale. The deadline is judged under the row's lock, after any wait for it, so a tick that crossed it while another worker or an admin save held the row is dropped rather than run late. The default is no deadline, which is the behaviour settings-declared schedules have always had. A dropped tick logsschedule_tick_droppedwith its lateness, once per worker, so it can be alerted on instead of vanishing.django_ox.stored.DatabaseScheduleSource, which dispatches the stored schedules alongside anySCHEDULESentries. Name it inOPTIONS["SCHEDULE_SOURCE"]and a row can be created, retimed and paused without a deploy or a restart. Rows are re-read when the change row moves, and in full everyOPTIONS["SCHEDULE_RECONCILE_INTERVAL"]seconds (default 60) as the backstop for a row written withoutdjango_ox.stored, so the steady state is one read of one row per dispatch pass. A row that no longer validates is skipped and logged asschedule_row_skipped; the others still run. When the rows cannot be read at all, the worker logsschedule_source_unavailable, with the database's own message and no traceback, and keeps dispatching the set it last read.- A schedule disabled, retimed or deleted after a worker read it does not
fire. The check runs inside the dispatch transaction, under the row's own
lock, because no polling interval is short enough to close that window. It
is taken with a locking read where the database has one, and on SQLite by
making the transaction a writer before it reads, because
select_for_update()there is a silent no-op. A refused tick commits nothing. django_ox.stored.create_schedule,update_scheduleanddelete_schedule, the supported programmatic write path. They validate, maintain the boundary and bump the change row. They take the fields a schedule's author owns and refuse the rest with aTypeErrornaming what they do take: the activation boundary, the count of writes to it and the two timestamps are theirs to maintain, and a misspelled field name is reported rather than set on the instance and silently not saved.create_schedulestill takesstart_time, which is a new schedule's boundary.save()does not callfull_clean(), so aclean()method alone would validate what the admin submits and nothing thatobjects.create()writes. A valid row written that way still runs, with its boundary moved to the read that found it and logged asschedule_boundary_healed.- A Django admin for stored schedules, the first place the package lets a
person author a row rather than act on one a worker wrote. The task field
is a choice drawn from the registry, so it cannot
express a task the code has not exposed, and the same membership check runs
again on the model for the write paths that build no form. A registry
entry's own
permissioncomes back as an error on that field, with the submission intact, rather than as a bare 403 that discards it; the permission itself is enforced indjango_ox.storedas before, on every write path. Saving routes throughdjango_ox.stored, so a schedule retimed in the admin gets its activation boundary moved rather than keeping one set for its old timing. Actions enable, disable, and run a schedule once immediately; a manual run writes no tick row, so the next scheduled tick still fires. It is also the one way to run a paused schedule: it ignores bothenabledandend_time, and says how many of the schedules it ran were disabled or already ended, and how many it was refused by a registry entry's own permission. The changelist carriesend_time, so a row that has ended is visible where the selection is made. The backend a manual run enqueues through is found the way the worker finds it: the classOPTIONS["SCHEDULE_SOURCE"]names is built and asked whether it answersschedules(), so any source the worker would dispatch from is found, whatever it is called and whatever it inherits, and a subclass is the class the run builds its schedule with. The changelist and the add page say so when no backend names a source at all, because until one does, a schedule saved there is stored and never dispatched, with nothing else to say so: no error, no log, no system check, and a manual run that enqueues anyway. manage.py ox_import_beat_schedules, which reads adjango-celery-beatschedule table and prints the django-ox equivalents. It writes nothing, names what it could not translate and why, and says which timing will differ.django_ox.E008reports django-ox models routed to more than one database. A task row and its tick row commit together, which is what makes a due tick enqueue once, so they must share a database. Routing the app to a single non-default database is supported.django_ox.E009anddjango_ox.W002report two configured schedule names that differ only by case. Tick identity is decided by the column's collation, and MySQL's default folds case, so the two share one key: their ticks collide and one schedule stops running with nothing raised. Refused on MySQL, warned about elsewhere.django_ox.W001reportsUSE_TZ = Falsetogether with aTIME_ZONEthat puts the clock back once a year. A tick's time is stored as a wall clock there, so one label covers both passes of the repeated hour: an interval schedule loses about half its runs for the length of it, and a cron schedule inside it fires once rather than twice. The schedules page states the effect.schedule_lock_unavailable: a lock the database gave up waiting for, MySQL's lock-wait timeout or a deadlock it resolved against this worker, SQLite's busy timeout, PostgreSQL'slock_timeout, is logged once without a traceback, and the schedule is skipped this pass; its tick fires on a later pass if still unclaimed. In 1.1.0 the same timeout on a settings-declared schedule's tick ended the whole poll pass asworker_poll_failed, with a traceback, and the claim did not run that pass.
Fixed
- A settings-declared schedule anchors once, however many workers first see it and whatever tick each of them holds. Two workers whose first passes fell either side of a minute boundary, or whose clocks differed by one, each wrote an anchor: their rows had different tick times, so the unique constraint serialised neither, and neither read saw the other's uncommitted row. The later instant was then claimed with no task, and the run it was for never happened. A pass that reads no history now takes a per-schedule latch inside its transaction before it decides, so the second worker waits, sees the anchor, and fires. The latch is a tick row at an instant no trigger produces, 1900-01-01, written and deleted inside the transaction, so nothing ever reads it. Present in 1.1.0.
- The tick read at the start of a dispatch pass fits the database's
parameter limit. It names every schedule in one
INlist, and SQLite before 3.32.0 refuses a statement with more than 999 parameters, so a deployment with that many schedules on SQLite 3.31 dispatched nothing, every pass. The keys are now read in slices of what the connection allows; PostgreSQL, MySQL and current SQLite declare no limit and read them in one statement as before. Present in 1.1.0. - An enqueue that fails with an integrity error is reported as
schedule_dispatch_error. In 1.1.0 every integrity error inside the dispatch block was read as a lost race with another worker, so an enqueue that kept failing was retried silently on every tick. The two are now told apart by whether this pass's own tick row had gone in. - A
transaction.on_commitcallback that raises after a dispatch commits is logged asschedule_dispatch_callback_failed, with the task id, and the dispatch is counted, since the task exists and the tick is recorded, whatever the callback raised: a database error from a callback's own statement is the callback's too, not a lost race or a fault that ends the pass. In 1.1.0 the exception left the dispatch pass with the task committed and, unless it was a database error, leftrun()too and stopped the worker. The stored-schedules page states the lock order atask_enqueuedreceiver must respect: it runs inside the dispatch transaction, under the schedule row's lock.
Changed
- A database error inside the dispatch pass ends the pass, logged once as
schedule_dispatch_failed; a connection that is no longer usable is dropped, and the claim still runs on a fresh one. In 1.1.0 the same error ended the whole poll pass asworker_poll_failedand the claim waited for the next one. Anything else one schedule raises is logged against that schedule asschedule_dispatch_errorand the rest of the pass continues. - Schedule dispatch runs on every pass whether or not a schedule is configured, because a source that reads the database can gain one at any time. With nothing configured it returns on a list check, before any query.
- The recurring-tasks page states what the tick log guarantees. It deduplicates the enqueue: a unique constraint on (schedule name, tick time) stops two workers enqueueing the same tick. It does not make a task run exactly once, and it does not enqueue every cron occurrence, because only the latest missed tick is enqueued after an outage and a new schedule never fires for a time before it existed. The README, the production page and the agent-facing files say the same.
1.1.0 - 2026-09-11
Two migrations ship with this release. 0005_dequeue_index rebuilds the
index the claim reads and adds a second one; 0006_lease_expiry adds a
nullable column and an index on it. Migrate before rolling any process that
imports django-ox 1.1.0, web processes included: enqueue() writes the column
0006 adds.
On PostgreSQL, CREATE INDEX takes a lock that blocks enqueues and claims for
the duration; MySQL 8 builds a secondary index online. On a large PostgreSQL
task table, build the indexes by hand and fake the migrations, in an order
that keeps the claim indexed throughout. Build the replacement
ox_dequeue_idx under a temporary name, swap it in, then add the other two:
CREATE INDEX CONCURRENTLY ox_dequeue_idx_new
ON django_ox_oxtask (status, priority DESC, enqueued_at);
DROP INDEX CONCURRENTLY ox_dequeue_idx;
ALTER INDEX ox_dequeue_idx_new RENAME TO ox_dequeue_idx;
CREATE INDEX CONCURRENTLY ox_dequeue_queue_idx
ON django_ox_oxtask (status, queue_name, priority DESC, enqueued_at);
ALTER TABLE django_ox_oxtask
ADD COLUMN lease_expires_at timestamp with time zone NULL;
CREATE INDEX CONCURRENTLY ox_reaper_expiry_idx
ON django_ox_oxtask (status, lease_expires_at);
then migrate --fake django_ox 0006. The ADD COLUMN is nullable with no
default, which is a catalogue change. sqlmigrate django_ox 0005 and
sqlmigrate django_ox 0006 print the same statements without
CONCURRENTLY, for checking against your settings.
Added
lease_expires_aton the task row: when the lease stops being valid, written by the worker that took it and refreshed on every renewal. Every reaper judges that column instead of deriving a deadline from its ownLOCK_TIMEOUT, so the setting can be changed in a rolling deploy.
A row claimed before the column existed has it empty, and the reaper keeps
comparing locked_at against its own timeout for those. The first renewal
after the upgrade fills it in, so a fleet converges lease by lease with
nothing for an operator to run. An expiry older than the row's own
locked_at was left there by a worker that does not know the column and
counts as absent too, so with LOCK_TIMEOUT unchanged across the rollout
a 1.0.0 worker that re-claims a row is judged on locked_at.
- django_ox.actions.expire_lease(result_id) expires a RUNNING task's lease so
the next reaper pass reclaims it. The row carries its own deadline now, so a
lease granted with a timeout that turned out to be wrong outlives the
configuration that granted it; this is how to shorten one. It does not stop
the task.
- lease_expires_at appears in the admin's Lease fieldset.
Fixed
- A worker that renews its lease late, but renews, keeps its task. The reaper checks the expiry in the reclaim itself, against the same cutoff it selected with, so only a lease still stale at the moment of the write is taken back.
- The lease renewal thread survives a dropped connection. It discards the
connection, reconnects on the next interval, and keeps
locked_atmoving on every row the worker holds. - The poll loop survives a database error. The pass is abandoned, logged as
worker_poll_failed, and retried on the next one with a fresh connection, so one blip is not a worker death and cannot reach the supervisor's restart cap. - The timeout watchdog handles each stuck attempt on its own. A failure
recording one is logged as
watchdog_errorand every other armed attempt keeps its deadline, and the recycle that frees the worker happens whether or not the record was written. - A worker whose timeout backstop gives up on a thread recycles on whether that thread is still running, not on whether its own outcome write landed. The pool slot is freed and the drain does not wait on the thread.
- The drain waits for healthy work and stops waiting for abandoned work. It counts threads still inside the attempt that was abandoned, rather than pool threads that happen to be alive, and it observes a backstop that fires part way through an ordinary drain.
- The drain is safe to run while a second attempt goes stuck: the stuck set is written under the same lock the drain reads it under.
- Every statement in the claim protocol runs on the database the worker writes to, including the raw PostgreSQL claim, so a read replica cannot answer a question that decides who holds a task.
enqueue_manyopens its transaction on the connectionOxTaskroutes to, so its all-or-nothing guarantee holds under a database router.- A worker subclass that overrides
claim_filter_q()alone takes the claim path that applies it, on every database. It logsclaim_filter_sql_missingonce to say it gave up the single-statement PostgreSQL claim to do so. - The PostgreSQL claim, the renewal and the reaper stamp and judge the lease
on one clock: the database's with
USE_TZon, the worker's with it off. - The reaper requeues abandoned rows in one UPDATE per pass, up to
reap_batchrows at a time, and retires rows whose attempts are spent in the same bounded batches. Its cost per pass is bounded whatever the size of the stuck set. - The reclaim record names only tasks the pass reclaimed. A pass
whose stuck set changed while it ran reports a
countand names nobody. - The claim reads its candidate out of an index, in order, for a worker that
names one queue, several, or none. Two index shapes ship and the planner
picks per query. Migration
0005_dequeue_indexbuilds them. - Dispatching schedules reads only the ticks it is asking about: the tick log is bounded to the oldest due tick across the configured schedules, which the unique index can seek to, and the read is pinned to the database the worker writes to.
- The tick row is written before its task is enqueued, so on every tick
exactly one worker enqueues and the others announce nothing. A tick already
in the log is not dispatched again whatever a fast-clocked worker recorded
after it, and
task_enqueuedfires only for a task that exists. - A signal receiver that raises is not charged to the task. A
task_startedreceiver's exception does not spend an attempt, and atask_enqueuedreceiver's exception does not surface fromenqueue()over a task that is already committed. - A task run through
run_once()keeps its lease renewed for the duration, the same as a task on the pool. KeyboardInterruptandSystemExitreach the caller when a task runs throughrun_once(). On the worker pool they are recorded as a failed attempt, so one task callingsys.exit()cannot stop a fleet.- The retry delay stays in range for any
MAX_ATTEMPTS. The doubling is capped before the multiplication, so the failure path never raises on the arithmetic. LOCK_TIMEOUT,BACKOFF_INITIALandBACKOFF_MAXare validated atmanage.py checkasdjango_ox.E010, by the same rule as the timeout options: a positive, finite number of seconds.- Each stored traceback is capped at 16,384 bytes, marker included, with both ends kept. One failure cannot write an unbounded string onto its own row.
django_ox.remaining()is measured on the monotonic clock the timeout is enforced with, so a clock correction cannot put the two on different sides of the same instant.django_ox.deadline()still answers with a wall-clock time.manage.py ox_workerstarts on Windows. Stop signals are built from the ones the platform has;--processesabove 1 still refuses off POSIX with its usual message.
Changed
- A recycling worker now bounds how long it waits for its other in-flight
tasks, at
LOCK_TIMEOUT. Past that it stops renewing their leases and exits, and the reaper requeues them. This can end a task that was going to finish: it is recorded as a lost lease and retried, and one on its last attempt reaches LOST. The bound is what makes a recycle certain on a process that has stopped trusting one of its threads. MAX_ATTEMPTSand theattemptlog key count claims, not invocations. The behaviour is unchanged; the reference table now says so.- Five log events that only fire while something is wrong are documented:
worker_poll_failed,watchdog_error,task_stuck_unrecorded,worker_drain_abandonedandclaim_filter_sql_missing. task_reclaimedcarriesheld_by, the worker that stopped refreshing the lock.worker_idon the same record is the reaper that noticed.- SQLite guidance: run one worker and give it threads with
--concurrency, rather than several worker processes.SKIP LOCKEDis what lets workers step past each other's rows, and a database without it hands out the head of the queue one worker at a time. task_reclaimedis still one record per task on the ordinary path. When the stuck set changes underneath a reap pass (a lease renewed, or one more expired), that pass emits a single record carryingcountand notask_id, rather than naming tasks it cannot vouch for.- The production page says which databases give the lease one shared clock.
SQLite computes
Now()inside the process that runs the statement, so every worker uses its own clock there whateverUSE_TZsays; run SQLite on one host. WithUSE_TZoff the lease is on the worker's clock on every database, so keep worker clocks within two thirds ofLOCK_TIMEOUTof each other, the timeout less the renewal interval.
1.0.0 - 2026-09-05
This release marks django-ox as production ready. The public API is stable from here: anything documented keeps working until a 2.0, and breaking changes get a major version. The package runs on Django 5.2 LTS through 6.1, on SQLite, PostgreSQL and MySQL, and the suite covers all of them.
Added
- Django 5.2 LTS support. Django ships the Tasks framework in core from 6.0;
on 5.2 the same framework comes from the
django-tasksbackport, and a newbackportextra pulls it in:pip install "django-ox[backport]". Every import of the framework now goes throughdjango_ox.compat, which picks core or the backport at runtime, so nothing else in the package changed. CI runs the whole suite on 5.2 against the backport, on SQLite, PostgreSQL and MySQL. Python 3.14 is not in the 5.2 legs, since Django 5.2 does not support it. Django>=5.2replacesDjango>=6.0as the declared dependency, andFramework :: Django :: 5.2joins the classifiers.
Changed
finished_atis stamped from the process clock instead of being computed by the database and read back, so an outcome write is one statement. The lease is untouched:locked_atand the reaper's cutoff stay on the database clock, which is where a lease is judged.- The
backportextra requiresdjango-tasks>=0.12. Earlier backport releases change the framework API in ways django-ox does not support: 0.9.0 has noenqueue_on_commitonTask, 0.10.0 rejects the result statuses django-ox stores, and 0.11.0 adds an abstractsave_metadatathat the backend does not implement. CI now installs the floor with==and asserts the resolved version. Development Status :: 5 - Production/Stablereplaces the beta classifier.
0.4.0 - 2026-09-01
Added
WORKER_CLASSin a backend'sOPTIONS: the dotted path of theWorkersubclassox_workerruns. Resolved from settings rather than from the command line, so a worker chosen here is the worker in every child process underox_worker --processes N, not only at--processes 1.Worker.claim_filter_q()andWorker.claim_filter_sql(), two hooks a subclass overrides to narrow what it may claim. The condition is applied inside the candidate select on all three claim paths, ahead of its ordering and its limit, so a narrowed worker still claims runnable work behind rows it declines.claim_filter_q()covers the two paths that build a queryset andclaim_filter_sql()the PostgreSQL claim, which builds its own SQL, so implement both. The defaults areNoneand an empty fragment, so the emitted SQL is unchanged without them.
0.3.1 - 2026-08-24
Fixed
ox_worker --processes Nexits 0 when a stop signal arrives while a worker process is still starting up. A worker cannot act on a signal until it has installed its handler, and that is after Django has been imported. A signal landing before then killed the worker outright, and the supervisor reported that worker's 143 as its own exit code. A unit onRestart=on-failurereads that as a fault and starts the service again. The worker had claimed no work, so there is nothing to report: the supervisor now logsworker_process_stopped_earlyand exits 0.
Changed
- A worker enforces
TASK_TIMEOUTon a sync task through the grace backstop alone while a coverage tool or a debugger is watching the thread the attempt runs on. Two things count: a trace function (sys.settrace, whichcoverage runinstalls before Python 3.14, and whichpdband most debuggers install), and a registeredsys.monitoringtool with events enabled (whichcoverage runuses from Python 3.14 on). Nothing is raised inside the task. A task that returns withinTASK_TIMEOUTplusTASK_TIMEOUT_GRACEis recorded as whatever it did, however long it ran, with notask_timed_outevent; one still running then is recorded as failed and recycles the worker. An async task is cancelled at its deadline either way, and a profile hook (sys.setprofile) is not consulted. - The worker logs
timeouts_backstop_onlyonce under such a tool, on the first attempt it registers rather than at startup. That event now carriesreason(interpreterortracing_tool) and, for a tool,tracer:sys.settrace, orsys.monitoring (NAME).
0.3.0 - 2026-08-24
Upgrading
- Run
python manage.py migrate django_ox, and roll every process, web and worker, to 0.3.0 before the first discard. This release adds the DISCARDED status: a 0.2.1 process that reads a DISCARDED row raisesValueErrorfromget_result()andrefresh(), and 0.2.1'sox_prunecannot delete such rows. Rolling back with DISCARDED rows present keeps that crash until the rows are removed by hand (DELETE FROM django_ox_oxtask WHERE status = 'DISCARDED'); reversing the migration does not remove them.
Added
TASK_TIMEOUT,TASK_TIMEOUTSandTASK_TIMEOUT_GRACEbackend options, a limit on how long one attempt may run. At the deadline the worker raisesdjango_ox.exceptions.TaskTimeoutinside the task, on the task's own thread, sofinallyblocks run and an opentransaction.atomic()rolls back; an async task is cancelled inside its event loop instead. The attempt is recorded as failed with theTaskTimeouterror and retried on the usual backoff, or marked FAILED when attempts are spent. A thread that has not stoppedTASK_TIMEOUT_GRACEseconds later (default 30) is treated as stuck: the worker records the attempt as failed, moves the lease number so the thread can write nothing to the row, stops claiming, drains its other tasks and exits with code 75, which--processesrestarts without counting it against the restart cap.TASK_TIMEOUTSmaps a queue name to its own value, and a key that is not inQUEUESfailsmanage.py checkasdjango_ox.E005; a bad value fails it asdjango_ox.E004. Off by default.django_ox.deadline()anddjango_ox.remaining()read the attempt's deadline from inside a task.TaskTimeoutsubclassesTimeoutError. New log events:task_timed_out,task_stuck,worker_recycling,worker_process_recycledandtimeouts_backstop_only.ox_worker --processes N. Above 1, the command supervises N copies of itself, each a full worker with its own database connections, lease renewal, reaper and--concurrencythread pool, so--processes 2 --concurrency 4runs eight tasks at once. Every worker id ends in its slot number. SIGTERM, SIGINT or SIGHUP to the supervisor is forwarded once and the supervisor exits 0 when every worker drained or recycled, or with the first other non-zero code; a second signal is the force-exit, and a worker that has not exited five seconds later is SIGKILLed. A worker process that dies is restarted withworker_process_restartedat WARNING, after one second and then with a doubling delay up to 30 seconds; a worker that recycled itself, exit code 75, comes back after one second underworker_process_recycled, outside the death count and the backoff. More than five deaths of one slot in a minute stops the supervisor with exit code 1 andsupervisor_restart_capat ERROR. A worker whose supervisor dies drains and exits (worker_orphaned). The children are started the way the supervisor was,manage.pyby absolute path orpython -m django, with--settingsand--pythonpathpassed on, so the command works from any working directory.--processes 1, the default, is the worker as before. POSIX only.- A Prometheus endpoint.
path("ox/", include("django_ox.urls"))mountsGET /ox/metrics, which renders thedjango_ox.statsnumbers as gauges in the Prometheus text format (OpenMetrics on request), from the standard library alone. The metric names aredjango_ox_tasks{queue,status},django_ox_ready_tasks,django_ox_oldest_ready_age_seconds,django_ox_last_claim_age_seconds,django_ox_throughput_per_minuteanddjango_ox_failure_rate, and they are public API from this release. The view has no authentication of its own.django_ox.metrics.collector()returns a collector for aprometheus_clientregistry when that package is installed; it is not a dependency. django_ox.actions.retry(result_id)anddjango_ox.actions.discard(result_id). A retry puts a FAILED or LOST task back to READY for one more attempt, keeping its attempt count, worker ids and every traceback, and clearing the backoff so it is eligible at once. A discard closes a READY, FAILED or LOST task without running it. Each is one compare-and-set on the row's status and lease number, so two retries of one row requeue it once, a discard that races a claim loses to it, and a LOST row's missing worker cannot write over its retry. Neither touches a RUNNING task.retry_many(selection)anddiscard_many(selection)take a queryset or a list of ids and make the same move in one conditional UPDATE per thousand rows inside one transaction, returning(changed, skipped).- The task table in the Django admin, registered only when
django.contrib.adminis installed: a list with status and queue filters and search by id or path, a read-only detail page with every attempt's traceback, and Retry selected tasks and Discard selected tasks actions that report how many rows moved and how many were skipped. The actions callretry_manyanddiscard_many, so a select-across of any size is a few statements in one transaction. The admin does not add, edit or delete rows. django_ox.bulk.enqueue_many(task, calls), the bulk form ofenqueue().callsis a list of(args, kwargs)pairs; the rows are written with one INSERT per 1,000 inside one transaction and theTaskResultlist comes back in input order. The task is validated and every argument serialised before the first write, so a rejected call inserts nothing. Each row is built by the same code asenqueue(), so workers see no difference.OxTask.Status.DISCARDED, a sixth value in django-ox's own status column. It reads asFAILEDthroughdjango.tasksandis_finishedis true for it.queue_stats()reports it in adiscardedcolumn, andox_prunedeletes discarded rows with successful ones.
Fixed
manage.py checkruns the django-ox checks,django_ox.E001toE005, in a project that importsdjango.tasksnowhere else. Django registers its tasks check when that module is first imported, and a project without the admin or a task module on its import path reachedcheckwithout it, so every django-ox check passed silently. The worker's own startup check was unaffected.
Changed
- A migration ships with this release. Run
python manage.py migrate django_oxwhen you upgrade. It adds the new status choice. QueueStatshas a sixth field,discarded, keyword-defaulted likelost.- A queued task can now be discarded and a failed or lost one retried, and
every attempt can be bounded with
TASK_TIMEOUT. - A worker that is already draining, because it is recycling, treats the operator's first signal as the drain it is doing rather than as the force-exit; the second signal is still the force-exit.
0.2.1 - 2026-08-20
Fixed
- A task that succeeds after its lease was lost drops the reaper's
lost-lease record from
errors. That record says the outcome was never observed, and the success write is that observation, so nothing readingresult.errorsis handed an exception nobody raised. Every earlier attempt's traceback stays on the row. A failure resolving the same way already dropped it, and the two now agree. A task that is still LOST keeps the record: it is the only thing on the row that says why the result reads as failed.
Added
- Built distributions are checked for every migration before release.
0.2.0 - 2026-08-20
Fixed
- A worker whose task had been taken back by the reaper could still write its own outcome over the row, so a task that had already finished could be moved back to READY and run a second time after its result had been reported. Every claim now stamps the row with a lease number, and every finish write carries that number in its WHERE clause, so a write from a worker that no longer holds the task matches nothing and is dropped instead of applied. No completion is signalled for a dropped write.
- A lock that ages out with no attempts left is recorded as LOST rather than
FAILED. LOST says the worker stopped reporting and the outcome was never
seen, and nothing more. The
TaskAbandonedrecord the reaper leaves inerrorsis the lost lease, not a cause of failure. - Lock timestamps are written and compared using the database server's clock
rather than each worker's own, so two hosts with drifting clocks no longer
produce false reclaims. This applies when
USE_TZis on. WithUSE_TZoff the worker's clock is used instead, because the database's clock does not always match what these columns hold: SQLite's is UTC while the columns carry naive local time, and reading one against the other would makeox_prune --older-thantreat rows that finished seconds ago as hours old. - On databases without
SELECT ... FOR UPDATE SKIP LOCKED, which includes SQLite, a claim read its row back in a second statement and could come away holding a lease granted to a different worker, if the reaper reclaimed the row in the gap between the two. The read is now pinned to the lease the claim was granted, so a worker that lost the row inside that gap comes back with nothing rather than with someone else's lease.
Added
- Lease renewal. A worker refreshes the lock on the tasks it is running,
one statement per interval however many are in flight, and keeps doing so
through a graceful drain. A long task on a healthy worker is no longer
reclaimed while it is still running. [Editorial note added in 1.4.0:
releases 0.2.0 to 1.3.1 were still affected by renewal starvation. See the
1.4.0 entry for the fix.]
LOCK_TIMEOUTnow bounds how long a worker may go unresponsive, not how long a task may take. The renewal interval isLOCK_TIMEOUT / 3, overridable asrenew_intervalwhen embeddingWorkerdirectly. OxTask.Status.LOST, a fifth value in django-ox's own status column. It reads asFAILEDthroughdjango.tasks, which has four statuses and no fifth, andis_finishedis true for it, so callers waiting on a result still terminate. The row keeps the distinction:queue_stats()reports alostcolumn andox_prune --include-failedcovers it. If the worker holding a LOST task comes back and records a real outcome, that outcome replaces LOST; only that one execution can.task_lease_lostandlease_renew_failed, two WARNING log events. Both are documented on the Monitoring page.
Changed
- A migration ships with this release. Run
python manage.py migrate django_oxwhen you upgrade. It adds thelease_epochcolumn and the new status choice. task_reclaimednow reportsstatusasREADYorLOST, where it previously reportedREADYorFAILED.QueueStatshas a fifth field,lost. It is keyword-defaulted, so existing code that constructs one keeps working.
0.1.2 - 2026-08-18
The worker, the public API and the database schema are unchanged. This release updates the project description that appears on the package page, and the documentation that ships with it.
Changed
- README now leads with what the backend removes from a deployment: the queue lives in the database the application already runs, so there is no broker to provision, secure, upgrade or back up. The transactional guarantee follows it rather than opening.
Added
- Migration guidance now covers moving away from django-ox as well as to it: which behaviour carries over to a broker-backed backend, which does not, and how to keep the option open.
- Worked examples for routing a queue to its own worker, choosing a lock timeout for long tasks, overriding a schedule's queue and priority, verifying that a schedule is live, and running the worker in containers.
context7.json, so documentation indexers read the project description, the supported versions and the setup steps rather than inferring them.
0.1.1 - 2026-08-17
The worker, the public API and the database schema are unchanged. This release updates the packaging metadata and the project description that appears on the package page.
Changed
- Packaging metadata now carries a
DocumentationURL, so the documentation site is linked directly from the package page. - README now carries release and CI status badges, a link to the documentation site, and a scope statement: what the core covers, what is deliberately outside it, and which features belong to the commercial tier.
0.1.0 - 2026-08-16
Initial release.
Added
OxBackend, a database-backed backend for Django's Tasks framework (django.tasks, Django 6.0+). Tasks are stored in the application database; no broker required.- Transactional enqueue:
enqueue()is a single INSERT on the caller's connection, so a task enqueued insidetransaction.atomic()commits or rolls back with the business data. ox_workermanagement command: claims tasks withSELECT ... FOR UPDATE SKIP LOCKEDwhere supported (PostgreSQL, MySQL 8+) and an atomic compare-and-set UPDATE elsewhere (including SQLite). Configurable via--backend,--queues,--concurrency(thread pool),--interval, and--lock-timeout.- Retries with exponential backoff (
MAX_ATTEMPTS,BACKOFF_INITIAL,BACKOFF_MAX), keeping the full traceback of every attempt. - Reaper: tasks whose worker died are returned to the queue after
LOCK_TIMEOUTand count as a failed attempt. - Graceful drain: on SIGTERM/SIGINT the worker stops claiming, finishes in-flight tasks, then exits; a second signal forces an immediate exit.
- Priorities (-100 to 100, higher first) and deferred tasks (
run_after), with the correspondingsupports_*flags declared on the backend. - Result store:
get_result(),refresh(), and the async variants, with status, return value, and per-attempt errors readable from the database. ox_prunemanagement command: batched deletion of finished task rows (--older-than,--include-failed,--batch-size,--dry-run).django_ox.stats: read-only queue metrics as plain ORM queries, on every supported database: per-queue status counts, backlog depth and age, throughput and failure rate over a trailing window, and time since the last task claim.ox_healthmanagement command: exits non-zero with a one-line reason when the database is unreachable or a--max-backlog,--max-ageor--worker-timeoutthreshold is breached; built for cron alerting and container probes.- Structured logging: worker lifecycle events (claim, start, success,
retry, failure, reclaim, dispatch, shutdown) log to the
django_oxlogger with stable extra keys (event,task_id,queue,attempt,duration_ms, ...) for JSON log handlers. - Recurring tasks: cron schedules declared in the
TASKSsetting (SCHEDULESoption), dispatched by the workers themselves; a unique constraint on (schedule, tick) enqueues each due tick once across any number of workers. Five-field cron syntax plus@hourly-style shortcuts; misconfigured schedules fail at startup and inmanage.py check. On recovery after downtime, only the latest missed tick fires. - System check
django_ox.E003: a schedule name defined on more than one backend is rejected, at worker startup and inmanage.py check, because the tick log is keyed by schedule name alone and shared names would let the backends suppress each other's ticks. - Strict cron validation: expressions that can never fire and step values
larger than a field's range (such as
*/61in the minute field) are rejected at parse time rather than misfiring silently. Schedule dispatch holds under clock skew between workers: a tick row dated in the future cannot suppress ticks that are due.
Security
- A stored
task_pathmust resolve to adjango.tasksTask (a function registered with@task). A row naming any other importable callable is rejected as an un-runnable task instead of being executed, so the worker never invokes an arbitrary dotted path pulled from the table.SECURITY.mddocuments the full trust model, the JSON-only serialization, and the guidance to keep secrets out of task arguments. - An API stability and deprecation policy (
docs/stability.md) covers the public API surface, the pre-1.0 SemVer rule, the deprecation window, and the supported Python and Django matrix.