Health monitor: the Twilio check had silently stopped running
Shipped 2026-08-20
When health monitoring moved off the Dev-DB-Redis droplet to
.github/workflows/health-monitor.yml, thirteen of the fourteen checks carried
over cleanly. The fourteenth did not, and nothing said so.
Why it went quiet
twilioSms only ran when new Date().getHours() === 0. The droplet’s
0 * * * * crontab hit the top of every hour reliably, so the gate opened each
night. GitHub’s scheduler does not: it delays scheduled runs by 25–58 minutes
under load and drops the 00:07 slot outright. Across 200 consecutive runs,
not one landed in hour 00 UTC — so the check reported skipped every single
time.
Skips don’t alert. Only failures do. The check had been failing whenever it did
run — with Twilio credentials not configured — but once it stopped running at
all, the alerts stopped too, and a broken check became indistinguishable from a
healthy one.
The credentials were never actually missing. The check read TWILIO_PHONE and
NUMBER_VIVEK, names that existed only in the droplet’s local .env and never
in Infisical. The rest of the API uses TWILIO_PHONE_US / TWILIO_PHONE_UK
(see api/src/features/shared/sms/service), which were present all along. The
error listed all three credentials as missing when only the phone number was,
which sent the investigation looking in the wrong place.
What changed
Cadence is anchored to work done, not to the clock. The new rules live in
api/src/crons/health/domain/staleness.js as pure, unit-tested functions.
isDueForPeriodicRun gates on the last real attempt rather than the last
success — gating on success would re-run a failing check every hour and alert
every hour.
Skips now expire. escalateStaleSkips promotes any check whose last real
verdict is more than 48 hours old to a failure. A check that has stopped running
can no longer masquerade as one that is passing.
Alerting moved to the crons Slack channel. mail.js is deleted. The Mandrill
report only ever went to TO_EMAIL, which was unset in production, so it fell
back to a hardcoded personal address — nobody on the team had seen a health
report from this pipeline. Failures now post a formatted summary naming every
failing check and its reason.
A daily all-clear posts when everything passes. Failure-only alerting cannot distinguish a healthy platform from a dead monitor. One heartbeat every 24 hours makes the silence between them mean something.
Twilio is checked without sending an SMS. The check reads the account status and the last 24 hours of message logs instead of texting a hardcoded number. It still catches revoked credentials, a suspended account, Twilio being unreachable and a run of undelivered messages; it no longer proves end-to-end delivery to a handset. In exchange it costs nothing, wakes nobody, and runs every hour instead of in a window the scheduler kept missing.
op_health_checks housekeeping
The table had grown past 124,000 rows with nothing pruning it, and no code anywhere reads it.
- 90-day retention, swept in 5,000-row batches at the end of each run rather
than one large
DELETE, so the backlog drains without holding locks. Configurable viaHEALTH_CHECK_RETENTION_DAYS. - The staleness lookup is bounded to 30 days, so it reads a date range instead of scanning the whole table every hour.
- The DigitalOcean checks stopped storing metrics blobs.
doAppPlatformwas recording every app’s id, URLs, timestamps and a live memory/CPU sample — ~1,465 characters a row, and 53% of everything stored in the table — for data nothing reads and that DigitalOcean’s own monitoring page already holds. Gathering it also cost two extra API calls per app per run. Notes are now 1,465 → 164 characters fordoAppPlatformand 432 → 136 fordoDroplet, and each run makes 16 fewer DigitalOcean API calls.
No index was added. It was measured — idx_created_at takes the hourly lookup
from 72.8ms to 12.3ms against 124k rows — but the retention sweep holds the table
near 34k rows, where the unindexed query costs about 20ms. The sweep is the fix;
an index would have been decoration.