Timer Service: Delivery Skew & Queue Health
Runbook for the Notifications Timer Service — the handler that republishes notifications deferred by the send-window scheduler.
Dashboards
- Notifications Timer Service
- Cloud Tasks console: deliver-push · deliver-email
- Errors: Bugsnag — Notification Services
Known-good baselines
Measured over a normal production week. Use these to judge whether what you are looking at is actually abnormal.
| Signal | Normal |
|---|---|
| Delivery skew, avg | ~180ms past deliver_at |
| Delivery skew, max | Usually under 10s; occasional spikes to 40–56s |
| Delivery skew, min | −20ms to −41ms (slightly early — this is normal) |
| Handler latency (APM) | ~200ms avg, ~1.3s max |
| Replicas (production) | 2 available, 0 unavailable |
| Queue depth | Grows through the day, drains as each user's 9am window opens |
republish_time is persistently negative by a few tens of milliseconds because Cloud Tasks
dispatches marginally ahead of the scheduled instant. Do not alert on min < 0 — it is true
essentially all the time. The alert threshold is 1 minute early — far outside anything
dispatch jitter can produce, so anything that trips it is a real scheduling defect.
Delivery skew is high
The SLI is actual_publish_time − deliver_at, p99 under 60s.
- Check whether deliveries are happening at all. The skew metric is emitted only on a successful publish, so a total outage shows up as the graph going flat/no-data rather than spiking. Look at "Deliveries published" and "Pub/Sub publish success rate" first. Note that recipients are spread across timezones, so multi-hour gaps with zero publishes are normal — a quiet stretch is only meaningful when the queue simultaneously has pending work.
- Check the queue depth panels. If depth is climbing on
deliver-pushordeliver-emailwhile deliveries are flat, dispatch is falling behind and the skew you see is only the tail of what is actually queued. - Check handler latency and the 5xx rate. If the handler is returning 5xx, Cloud Tasks is backing off exponentially, which inflates skew even after the handler recovers.
- Check replicas.
replicas_unavailable > 0means pods are failing readiness and the queue is being drained by fewer workers than intended.
If the handler is healthy and the queue is draining, the skew is upstream: the tasks were enqueued
with a deliver_at that had already passed, or the enqueue itself was delayed.
A notification was delivered early
Confirm the magnitude before treating this as an incident — see the warning above. Tens of
milliseconds early is expected and is not worth investigating. The monitor fires only past a full
minute early, which means either the enqueuing handler computed the wrong deliver_at, or the
send-window/jitter configuration for that notification is wrong. Check sending_window and jitter_window in the notification's config, and
the scheduled_at / deliver_at attributes carried on the republished message.
Early sends are user-visible and cannot be recalled, so escalate rather than waiting to see if it recurs.
Queue depth is growing
Both queues are designed to accumulate through the day and drain overnight as each recipient's send window opens, so growth alone is not a problem. What matters is growth that does not drain.
- Compare depth against the same time on previous days on the "Queue depth over time" panel.
- If depth is genuinely not draining, check "Task dispatch attempts" — a flat attempt count with a rising depth means Cloud Tasks is not dispatching (queue paused, or the handler is rejecting).
- Check whether the queue has been paused in the Cloud Tasks console.
- Check the handler 5xx rate: sustained 5xx means every task is being retried with backoff, which collapses effective dispatch throughput.
Handler is returning errors
5xx (publish_error, read_error) — transient by contract; Cloud Tasks will retry with
backoff, so nothing is lost yet. publish_error almost always means Pub/Sub is unhealthy or the
target topic is misconfigured. Check the topic exists and the service account can publish to it.
4xx (bad_request) — deliveries are being permanently dropped. Cloud Tasks will not retry.
Break down the "Bad requests by reason" panel:
reason | Cause |
|---|---|
empty_body | Task enqueued with no payload — bug in the enqueuing handler |
malformed_envelope | Payload is not a valid ScheduledDelivery proto — usually a schema mismatch after a deploy |
missing_topic | Envelope has no topic set — bug in the enqueuing handler |
All three are producer-side defects. The notifications that hit them are already lost; the priority is stopping the bleeding by fixing or rolling back the enqueuing handler.
Escalation
Post in #incidents-notifications. Early deliveries and permanently-dropped (4xx) deliveries are user-visible and irreversible — escalate those immediately rather than monitoring.