Operations: background jobs & schedules ======================================== tigerdistribute's background work runs on `Temporal `_: the ``temporalworker`` service executes workflows, and **Temporal Schedules** initiate the recurring ones. This page is for operators and developers; the user-facing behaviour of each job is described in the user guide. Recurring jobs -------------- Every active tenant has six schedules, named after its slug: .. list-table:: :header-rows: 1 * - Schedule - Workflow - Cadence * - ``relay-`` - ``OutboxRelayWorkflow`` - every minute * - ``notify-`` - ``NotificationDeliveryWorkflow`` - every minute * - ``monitoring-`` - ``LicenseMonitoringWorkflow`` - daily 02:00 UTC * - ``attainment-`` - ``AttainmentSweepWorkflow`` - daily 02:30 UTC * - ``payrun-`` - ``PaymentRunWorkflow`` - daily 04:00 UTC * - ``digest-`` - ``ComplianceDigestWorkflow`` - daily 07:00 UTC * The **relay** delivers pending outbox events to the tenant's webhook endpoints (see the user guide's *Integrations* chapter for delivery semantics). * The **notification delivery** run sends pending email notifications, with outbox-style exponential backoff (60 s base, 1 h cap) and dead-lettering after 5 attempts. * **Monitoring** sweeps licenses, E&O coverage, CE requirements and contract terms, raises compliance alerts, and finishes by re-deriving every producer's **payability** verdict (flips publish ``ProducerPayabilityChangedV1`` outbox events). * The **attainment sweep** recomputes attainment for every *active* target plan, then computes incentive awards and reconciles paid ones; drafts and closed plans are untouched. * The **payment run** settles every producer whose payment cadence is due today (weekly on Mondays, monthly on the 1st, quarterly on the quarter's first day; *immediate*-cadence producers are settled inline at finalize/approve and never swept), consolidating their payables into ``PayoutInstructedV1`` instructions. It runs after the attainment sweep so anything computed overnight is payable. * The **compliance digest** rolls up not-yet-announced alerts into the staff digest and per-producer emails — deliberately after the 02:00 monitoring sweep, early enough to land in the morning inbox. All schedules use the *skip* overlap policy: if a run is still going when the next tick fires, the tick is skipped rather than stacked. On-demand workflows (``RegistrySyncWorkflow``, ``TargetAttainmentWorkflow``, ``ProducerScoringWorkflow``, and ``BulkJobCommitWorkflow`` for bulk-job commits over 500 rows) are started by API actions or commands, not by schedules. How schedules are provisioned ----------------------------- Schedules live in the **Temporal server's own database**, not in the app — once created they survive app deploys, worker restarts and container rebuilds. You do *not* need to re-create them per deployment. ``manage.py create_tenant`` Ensures the new tenant's schedules as the last provisioning step. Best-effort: if Temporal is unreachable the tenant is still created and the command tells you to run ``setup_temporal_schedules`` later. Pass ``--no-schedules`` to skip (environments without Temporal). ``manage.py setup_temporal_schedules`` Idempotent sweep over all active tenants (``--tenant `` to limit). Use it to **bootstrap** a fresh environment (run once after ``migrate``), to repair drift, and — with ``--replace`` — to apply cadence changes made in ``tigerdistribute/temporal/schedules.py`` (delete + recreate). .. note:: A **newly added** schedule (such as ``payrun-``) is not created on existing tenants until this command runs. After deploying the payment-run feature, run ``setup_temporal_schedules --replace`` once to register ``payrun-*`` on every existing tenant. Inspecting and controlling -------------------------- The Temporal Web UI (local: ``http://localhost:58233``) shows every schedule, its next/last run and each run's history. From the CLI:: docker exec tigerdistribute_local_temporal_admin \ temporal schedule list --namespace tigerdistribute # Stop deliveries for one tenant (e.g. their maintenance window): docker exec tigerdistribute_local_temporal_admin \ temporal schedule toggle --schedule-id relay-acme \ --pause --reason "maintenance" --namespace tigerdistribute One-off runs, without waiting for the next tick:: python manage.py trigger_relay # outbox relay sweep python manage.py trigger_monitoring # license monitoring sweep (Or use the API actions, e.g. ``POST /api/outbox-events/relay/``.) Outbox retention ---------------- The outbox is append-only: the relay marks each event ``published`` or ``failed`` but never deletes it. The partial ``idx_outbox_due`` index keeps the sweep fast regardless of size, but the table still grows, so prune it periodically. The command is **manual and opt-in** — it is deliberately *not* on a schedule:: # Delete delivered (published) events older than 90 days, all tenants: python manage.py prune_outbox_events --retention-days 90 # Preview first (deletes nothing): python manage.py prune_outbox_events --retention-days 90 --dry-run ``--retention-days`` is required (no default), so a prune is never accidental. By default only ``published`` events are removed; pass ``--include-failed`` to also prune dead-lettered ``failed`` events. ``pending`` events are never touched at any age. Scope to one tenant with ``--tenant ``; deletes run in batches (``--batch-size``, default 5000). Confirm your audit/retention policy before wiring this into a cron/Temporal schedule. Local development notes ----------------------- * The ``temporalworker`` compose service runs under **watchfiles** and restarts automatically on any backend ``.py`` change — workflows and activities register at startup, so this is what makes code changes take effect. * Workflow inputs are dataclasses carrying ``tenant_id``; every activity re-establishes the RLS tenant context before touching the ORM. Schedules are therefore strictly per-tenant — pausing one tenant's schedule cannot affect another's. * Wiping the ``temporal-postgres`` volume (``docker compose down -v``) erases schedules along with the rest of Temporal's state — re-run ``setup_temporal_schedules`` after bringing the stack back up.