Scheduled jobs move off GitHub Actions, and start reporting what they did
Shipped 2026-09-22
Measured over eight days, GitHub’s scheduler ran the daily jobs 4 to 6 hours late, every day. Payment reminders scheduled for 06:00 UTC went out between 10:40 and 11:34. Backups scheduled for 02:00 ran around 07:13.
The hourly health workflow fared worse: 24 runs a day expected, 4 to 7 actual.
One cause, two symptoms. A roughly constant five-hour queue delay is survivable for a daily schedule — the run still lands before its next window, so it happens, just late. On an hourly schedule those same five hours swallow five windows, and GitHub drops a delayed scheduled run rather than stacking it. Which means the one workflow best placed to notice the problem is the one the problem silences.
Moving a timer, not a workload
The workflow never did the work. It made one authenticated GET to
/api/crons/run/<name>; the job itself executes on the API and records its own
outcome in op_cron_logs. So the fix only had to relocate the timer.
workers/cron-scheduler is a Cloudflare Worker doing exactly that. It lives at
the repo root rather than under api/, because the DigitalOcean app builds from
source_dir: api and anything there joins the API image.
The cutover is staged. GitHub still fires all seven real jobs. The Worker
fires only test_cron — a no-op that exits 0, which nothing in crons.yml
triggers — every 15 minutes. Sub-hourly on purpose: the failure being escaped is
dropped hourly slots, so a 15-minute cadence tests that directly and produces
96 samples a day instead of 24. Moving a job across means deleting its - cron:
from crons.yml and adding the same expression to wrangler.toml in the same
commit — both at once, or it fires twice or not at all.
The jobs were already reporting; nothing was listening
reconcile_parent_skills has ended with this for months:
logger.info({ event: 'reconcile_parent_skills', reconciled });That number went nowhere. The harness runs each script with
child_process.exec, which buffers the child’s stdout rather than
inheriting it, and reads that buffer only when the run failed. On success the
count was captured into a variable and dropped — it never reached the
container’s stdout, so it never reached Datadog either.
Scripts now print one CRON_SUMMARY {json} line. The harness scrapes it from
the buffer it already holds into the new op_cron_logs.summary column, and
re-emits it through the logger, where it does reach Datadog. Instrumented so far
where a count was already in hand: reconcile_parent_skills, payment_reminder
and health. The currency jobs emit nothing yet — their services return no
count — and NULL stays deliberately distinguishable from a reported zero.
Slack: one daily digest instead of a health-only all-clear
Failure alerts are unchanged. What changes is the only routine message. The daily all-clear reported the health checks and nothing else, so seven of the eight jobs were invisible unless they broke, and “no news” could not be told apart from “nothing ran”.
That slot now carries every job:
📋 Daily cron digest — 24h to 22:10 UTC
✅ Health checks 14/14 passing
✅ payment_reminder — 1 run, last 06:04, 12s — {"payments":38}
✅ reconcile_parent_skills — 1 run, last 04:00, 42s — {"reconciled":0}
✅ flyer_currency — 1 run, last 08:31, 1m10s — no summary
💤 no runs: trainer_currency{"reconciled":0} and no summary render differently on purpose. Zero is a
result; no summary is a gap. Collapsing them would recreate the ambiguity this
whole change exists to remove.
The digest also fires regardless of the health verdict now — a day with a failing check still wants one. The heartbeat is gone rather than left dead: the digest’s health line does the job it existed for, which was proving the monitor itself is alive.
A watchdog that can see lateness
dailyCron covered two of the nine jobs, by hardcoded cron_config_id, and
asked only whether they ran yesterday — which a five-hour delay passes
cleanly. That is why the drift went unnoticed for two months.
The new cronFreshness check reads every active row in op_cron_configs and
reports late as well as missing, measuring against each job’s own
schedule_pattern. Run against real data it reproduced the GitHub measurements
from op_cron_logs alone, knowing nothing about GitHub: flyer_currency late by
4h39m, payment_reminder by 4h21m.
It also immediately found a disagreement nobody had looked for.
reconcile_parent_skills settles on daily
op_cron_configs said hourly, per migration 0003’s “ongoing safety net”. The
workflow fired daily. Nothing compared them, so the job has run at one
twenty-fourth of its stated cadence since July.
Daily wins, on evidence. Migration 0001 stamps ip_address = 'reconcile:0001'
on every row the reconciliation writes, so its entire history is countable:
four rows, four members, all between 7 and 13 July, none since. It is a
safety net behind a race in the synchronous approval path; it caught four misses
in its first week and nothing in the ten after. That does not warrant 24 runs a
day. If the rate ever climbs, the new summary column now records
{"reconciled":n} per run, so the case for hourly can be made from data rather
than from a comment written before the job had ever run.