Timer Service
Notifications Timer Service exists at: https://github.com/joinhandshake/notifications/tree/main/services/timer
The timer service is the delivery half of the notification scheduling feature. When a
notification arrives outside the recipient's local send window, the producing handler
defers it instead of dropping it: it wraps the original Pub/Sub message in a ScheduledDelivery
envelope and enqueues a Cloud Task that fires at the (jittered) start of the next window. When the
task fires, Cloud Tasks calls POST /v1/deliver on this service, which republishes the wrapped
message verbatim to its original topic.
handler (outside send window)
└─> ScheduledDelivery envelope ──> Cloud Tasks queue (deliver-push / deliver-email)
│ fires at deliver_at
▼
POST /v1/deliver (timer service)
└─> Pub/Sub topic (original, byte-identical)
└─> existing DEP consumers, unchanged
Status code contract
Cloud Tasks retry behavior depends on the status the handler returns:
| Status | Meaning | Cloud Tasks behaviour |
|---|---|---|
| 2xx | Published successfully | Task acked |
| 4xx | Malformed or empty envelope | Permanent failure, no retry — the notification is lost |
| 5xx | Publish or body-read failure | Retried with backoff |
Metrics
All metrics are emitted from
internal/api/handlers/v1/deliver.go
and are prefixed notifications_timer_service. (the app name with dashes replaced by underscores).
Every metric carries env and service:notifications-timer-service.
| Metric | Type | Tags | Meaning |
|---|---|---|---|
timer.deliver.success | count | topic | Message republished |
timer.deliver.republish_time | distribution (ms) | topic | Delivery skew SLI — see below |
timer.deliver.publish_error | count | topic | Pub/Sub publish failed (5xx, retried) |
timer.deliver.bad_request | count | reason | 4xx — delivery permanently dropped |
timer.deliver.read_error | count | — | Request body unreadable (5xx, retried) |
reason on bad_request is one of empty_body, malformed_envelope, missing_topic. Every
failure path also notifies Bugsnag under the
Notification Services project.
The service is APM-traced, so trace.http.request.* is available for
service:notifications-timer-service, resource_name:post_/v1/deliver.
Delivery skew (primary SLI)
timer.deliver.republish_time measures actual_publish_time − deliver_at in milliseconds — how
far past its scheduled time a notification actually went out. It is the primary SLI for the
scheduling feature.
Target: p99 under 60 seconds, and no meaningfully early deliveries.
Two properties matter when reading or alerting on it:
- It is emitted only on a successful publish. A total publish outage makes the metric go no-data, not spike. Any skew alert therefore needs a companion signal — queue depth rising while successes sit at zero — to catch an outage.
- Negative values mean the notification went out early. A steady floor of roughly −20ms to −40ms is normal Cloud Tasks dispatch jitter and clock skew, not a defect. See the delivery skew runbook for the thresholds this implies.
republish_time is a Datadog distribution, but percentile aggregation is not currently enabled
on it, so p50:/p99: return no data and only avg:, max:, min:, sum: and count: work.
Enabling percentiles is a per-metric setting in Datadog and is not retroactive.
Deployment
Deployed via Shipit/ArgoCD as notifications-timer-service. Default port 5552. The
/deliver endpoint is unauthenticated at the application layer; OIDC is enforced at the ingress.