Skip to main content

Timer Service

Notifications Timer Service exists at: https://github.com/joinhandshake/notifications/tree/main/services/timer

The timer service is the delivery half of the notification scheduling feature. When a notification arrives outside the recipient's local send window, the producing handler defers it instead of dropping it: it wraps the original Pub/Sub message in a ScheduledDelivery envelope and enqueues a Cloud Task that fires at the (jittered) start of the next window. When the task fires, Cloud Tasks calls POST /v1/deliver on this service, which republishes the wrapped message verbatim to its original topic.

handler (outside send window)
└─> ScheduledDelivery envelope ──> Cloud Tasks queue (deliver-push / deliver-email)
│ fires at deliver_at
▼
POST /v1/deliver (timer service)
└─> Pub/Sub topic (original, byte-identical)
└─> existing DEP consumers, unchanged

Status code contract​

Cloud Tasks retry behavior depends on the status the handler returns:

StatusMeaningCloud Tasks behaviour
2xxPublished successfullyTask acked
4xxMalformed or empty envelopePermanent failure, no retry — the notification is lost
5xxPublish or body-read failureRetried with backoff

Metrics​

All metrics are emitted from internal/api/handlers/v1/deliver.go and are prefixed notifications_timer_service. (the app name with dashes replaced by underscores). Every metric carries env and service:notifications-timer-service.

MetricTypeTagsMeaning
timer.deliver.successcounttopicMessage republished
timer.deliver.republish_timedistribution (ms)topicDelivery skew SLI — see below
timer.deliver.publish_errorcounttopicPub/Sub publish failed (5xx, retried)
timer.deliver.bad_requestcountreason4xx — delivery permanently dropped
timer.deliver.read_errorcount—Request body unreadable (5xx, retried)

reason on bad_request is one of empty_body, malformed_envelope, missing_topic. Every failure path also notifies Bugsnag under the Notification Services project.

The service is APM-traced, so trace.http.request.* is available for service:notifications-timer-service, resource_name:post_/v1/deliver.

Delivery skew (primary SLI)​

timer.deliver.republish_time measures actual_publish_time − deliver_at in milliseconds — how far past its scheduled time a notification actually went out. It is the primary SLI for the scheduling feature.

Target: p99 under 60 seconds, and no meaningfully early deliveries.

Two properties matter when reading or alerting on it:

  • It is emitted only on a successful publish. A total publish outage makes the metric go no-data, not spike. Any skew alert therefore needs a companion signal — queue depth rising while successes sit at zero — to catch an outage.
  • Negative values mean the notification went out early. A steady floor of roughly −20ms to −40ms is normal Cloud Tasks dispatch jitter and clock skew, not a defect. See the delivery skew runbook for the thresholds this implies.
Percentile aggregation

republish_time is a Datadog distribution, but percentile aggregation is not currently enabled on it, so p50:/p99: return no data and only avg:, max:, min:, sum: and count: work. Enabling percentiles is a per-metric setting in Datadog and is not retroactive.

Deployment​

Deployed via Shipit/ArgoCD as notifications-timer-service. Default port 5552. The /deliver endpoint is unauthenticated at the application layer; OIDC is enforced at the ingress.

Dashboard​