Skip to content

feat(magento): durable:worker drains the database queues (#736) - #1005

Merged
gplanchat merged 15 commits into
mainfrom
feat/magento-durable-worker
Oct 9, 2026
Merged

gplanchat merged 15 commits into
mainfrom
feat/magento-durable-worker

Conversation

@gplanchat

Copy link
Copy Markdown
Owner

Refs #736. Epic #740, backend key database.

Stacked draft: this PR is based on #1003 (feat/magento-runtime-factory-sql), which sits on #1002 (feat/magento-sql-integration, a throwaway integration base). Merge order: #1003 into its base first, then this one repointed.

What changes

bin/magento durable:worker drains the table queues when resource/durable is declared.

  • No --role: one process serves the resume, timer and activity queues. --role=journal serves resumes and timers, --role=activity serves activities.
  • --role=nexus is refused: "The database backend has no Nexus role: Nexus operations need a Temporal cluster. Run durable:worker with --role=journal or --role=activity, or without --role to serve both."
  • SIGTERM and SIGINT end the loop between two messages.
  • DatabaseWorker::tick() takes a message, handles it under the execution lock (resume, timer) or the core processor's attempt claim (activity), and acknowledges it once handled.
  • A held execution, a resume that arrives before its outcome (DUR050), an activity attempt held elsewhere, a lock wait timeout, a deadlock or a lost connection give the message back with the new TableQueue::release() (one UPDATE, token-checked like ack()), delivered again after 0.5 s. The message is never acknowledged on those paths.
  • A body nobody can read, or a failure the handlers did not journal, is acknowledged and the run is journalled as failed, so that it does not wait for a message that is gone.
  • holds() on the execution lock is asked before each resume or timer handler. For activities the claim is held by ActivityMessageProcessor for the whole attempt; the worker has no hook inside the activity to ask again (see below).

The bench's minimal drain.php is replaced by magento/tests/worker.php, which runs the real command in its own process.

Tests

Bench, against MySQL 8.0 (magento/phpunit.xml.dist):

  • a real lock wait timeout (LOCK TABLES from a second connection, lock_wait_timeout = 1) leaves the resume in the queue, and it is handled once the lock is gone;
  • an early resume is re-queued and no failure is journalled;
  • a held execution lock re-queues, and the resume runs once it is released;
  • an unreadable body and an unknown workflow type are acknowledged, the second journalled as failed;
  • end to end with real worker processes: the first worker is kill -9ed inside the activity, the queue still holds the message, the server has freed the attempt claim, and a second worker finishes the run with one ActivityCompleted. The lease (600 s by default) is brought to its end with an UPDATE instead of being waited for;
  • SIGTERM makes the worker exit 0 with "stopped by a signal".

Root suite: the role refusals, with the messages above.

Not done here

  • The restart experiment of [Task] durable-magento: durable:worker on the Magento SQL backend #736 on the bench with the journal on a second MySQL server and the shop database holding 0 durable_* tables: the bench test above checks the second point (getTables('durable\_%') on the shop's connection is empty) and uses a second database on one server, not a second server.
  • A cap on how many times a message is given back: the table has no attempt counter.
  • holds() inside the activity, as described above.

@gplanchat
gplanchat changed the base branch from feat/magento-runtime-factory-sql to main October 9, 2026 12:01
@gplanchat
gplanchat marked this pull request as ready for review October 9, 2026 12:01
@gplanchat
gplanchat merged commit 7a57f3b into main Oct 9, 2026
47 checks passed
@gplanchat
gplanchat deleted the feat/magento-durable-worker branch October 9, 2026 12:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant