Skip to main content

Timer Service: Delivery Skew & Queue Health

Runbook for the Notifications Timer Service — the handler that republishes notifications deferred by the send-window scheduler.

Dashboards​

Known-good baselines​

Measured over a normal production week. Use these to judge whether what you are looking at is actually abnormal.

SignalNormal
Delivery skew, avg~180ms past deliver_at
Delivery skew, maxUsually under 10s; occasional spikes to 40–56s
Delivery skew, min−20ms to −41ms (slightly early — this is normal)
Handler latency (APM)~200ms avg, ~1.3s max
Replicas (production)2 available, 0 unavailable
Queue depthGrows through the day, drains as each user's 9am window opens
Early deliveries are normal at the millisecond scale

republish_time is persistently negative by a few tens of milliseconds because Cloud Tasks dispatches marginally ahead of the scheduled instant. Do not alert on min < 0 — it is true essentially all the time. The alert threshold is 1 minute early — far outside anything dispatch jitter can produce, so anything that trips it is a real scheduling defect.

Delivery skew is high​

The SLI is actual_publish_time − deliver_at, p99 under 60s.

  1. Check whether deliveries are happening at all. The skew metric is emitted only on a successful publish, so a total outage shows up as the graph going flat/no-data rather than spiking. Look at "Deliveries published" and "Pub/Sub publish success rate" first. Note that recipients are spread across timezones, so multi-hour gaps with zero publishes are normal — a quiet stretch is only meaningful when the queue simultaneously has pending work.
  2. Check the queue depth panels. If depth is climbing on deliver-push or deliver-email while deliveries are flat, dispatch is falling behind and the skew you see is only the tail of what is actually queued.
  3. Check handler latency and the 5xx rate. If the handler is returning 5xx, Cloud Tasks is backing off exponentially, which inflates skew even after the handler recovers.
  4. Check replicas. replicas_unavailable > 0 means pods are failing readiness and the queue is being drained by fewer workers than intended.

If the handler is healthy and the queue is draining, the skew is upstream: the tasks were enqueued with a deliver_at that had already passed, or the enqueue itself was delayed.

A notification was delivered early​

Confirm the magnitude before treating this as an incident — see the warning above. Tens of milliseconds early is expected and is not worth investigating. The monitor fires only past a full minute early, which means either the enqueuing handler computed the wrong deliver_at, or the send-window/jitter configuration for that notification is wrong. Check sending_window and jitter_window in the notification's config, and the scheduled_at / deliver_at attributes carried on the republished message.

Early sends are user-visible and cannot be recalled, so escalate rather than waiting to see if it recurs.

Queue depth is growing​

Both queues are designed to accumulate through the day and drain overnight as each recipient's send window opens, so growth alone is not a problem. What matters is growth that does not drain.

  1. Compare depth against the same time on previous days on the "Queue depth over time" panel.
  2. If depth is genuinely not draining, check "Task dispatch attempts" — a flat attempt count with a rising depth means Cloud Tasks is not dispatching (queue paused, or the handler is rejecting).
  3. Check whether the queue has been paused in the Cloud Tasks console.
  4. Check the handler 5xx rate: sustained 5xx means every task is being retried with backoff, which collapses effective dispatch throughput.

Handler is returning errors​

5xx (publish_error, read_error) — transient by contract; Cloud Tasks will retry with backoff, so nothing is lost yet. publish_error almost always means Pub/Sub is unhealthy or the target topic is misconfigured. Check the topic exists and the service account can publish to it.

4xx (bad_request) — deliveries are being permanently dropped. Cloud Tasks will not retry. Break down the "Bad requests by reason" panel:

reasonCause
empty_bodyTask enqueued with no payload — bug in the enqueuing handler
malformed_envelopePayload is not a valid ScheduledDelivery proto — usually a schema mismatch after a deploy
missing_topicEnvelope has no topic set — bug in the enqueuing handler

All three are producer-side defects. The notifications that hit them are already lost; the priority is stopping the bleeding by fixing or rolling back the enqueuing handler.

Escalation​

Post in #incidents-notifications. Early deliveries and permanently-dropped (4xx) deliveries are user-visible and irreversible — escalate those immediately rather than monitoring.