Monitor websites uptime using Cloudflare Workers
uptime docs operations notifications.md
8.7 kB
Markdown

Notifications #

Apply migrations before starting the updated API and scheduler:

pnpm db:migrate

Docker Compose's migration service does this automatically when deploying rebuilt images. Rebuild the API, scheduler, migration, and web images from this checkout; the existing published latest images do not contain local changes. Cloudflare Workers do not need to change for notifications.

Set up a destination #

Open Notifications in the admin panel. Add a name and choose a provider. Multiple destinations of the same provider are supported. Save, then use Test to check delivery.

Telegram #

  1. Open BotFather in Telegram and create a bot with /newbot.
  2. Copy the bot token into the form.
  3. Start a private conversation with your bot, or add it to your group. Send a message (a command addressed to the bot works in groups with privacy mode enabled).
  4. Retrieve the bot's getUpdates response and copy message.chat.id into Chat ID. Group IDs can be negative. A public channel username such as @your_channel is also accepted; the bot must have permission to post there.
  5. Save and send a test notification.

Keep the token private. getUpdates will not work while another integration has an active webhook for this bot; a dedicated notification bot avoids that conflict. Bots cannot initiate a private conversation until the recipient has started them.

The provider sends plain text using sendMessage, with link previews disabled.

Discord #

  1. Open your server's Server Settings → Integrations → Webhooks (or the channel's webhook settings).
  2. Create a webhook, choose its channel, and copy its URL.
  3. Paste the standard https://discord.com/api/webhooks/... URL, save, then send a test notification.

The URL grants permission to post in that channel. Treat it as a password. The provider uses Execute Webhook with wait=true to confirm delivery and disables automatic mentions, including @everyone.

Resend #

Create a Resend API key, verify the sending domain, then enter the API key, From email, and one or more To email addresses. Recipients may be separated by commas or newlines. Subject is optional. The service sends the shared notification text as a plain text email.

Gotify #

Enter the Gotify server URL, an application token created in Gotify, and a priority from 0–10 (default 8). The server URL may be HTTP or HTTPS for self-hosted installations. The provider posts a Gotify message with the notification text.

Generic webhook #

Enter an HTTP(S) webhook URL and, optionally, a bearer token. The provider sends a JSON POST body containing kind, monitorName, monitorUrl, occurredAt, outageStartedAt, and text. If a bearer token is configured, it is sent as Authorization: Bearer ….

SMTP #

Enter the SMTP host, port, security mode, sender, and recipients. Security defaults to STARTTLS; TLS and None are also available. Username and password are optional for relays that do not authenticate. None sends SMTP credentials and notification content in plaintext and should only be used on a trusted private network. Recipients accept comma or newline separated addresses, and subject is optional.

SMTP sends to recipients individually. If some recipients are accepted and a later recipient fails, the delivery is marked failed but is not automatically retried, because retrying could duplicate messages for recipients that already received them.

Home Assistant #

Enter the Home Assistant server URL, a long-lived access token, and a service slug such as mobile_app_phone. The service defaults to notify when left blank. The provider calls Home Assistant's notification service with the shared notification text.

Monitor rules #

In Add monitor or Edit monitor, choose any number of notification services, including none.

  • Failed checks before outage: 1–100 consecutive scheduled rounds, default 3.
  • Healthy checks before recovery: 1–100 consecutive scheduled rounds, default 2.
  • Repeat notifications while down: off by default. Enable and choose 1–10,080 minutes (up to seven days).

A round includes all regional checks for one scheduled monitor interval. It counts as failed if any returned regional result fails, including a target timeout or HTTP failure. It counts as healthy only when every expected regional check returns success. A round with missing results and no reported failure is unknown: it resets both streaks and does not declare an outage or recovery. Scheduler-to-Worker transport failures alone never declare the target down.

One outage message is queued when the failure threshold is reached. One recovery message is queued when the healthy threshold is reached. Reminders continue at the configured interval while the confirmed outage remains open, including while recovery is not yet confirmed. Missed reminders are collapsed into one; they are not replayed as a burst. With repeats disabled, a prolonged outage sends only the initial outage message and eventual recovery message.

These rules govern notification incidents. Existing instantaneous regional status and historical uptime percentages continue to describe the underlying observations.

Changes to the target, region set, thresholds, repeat policy, selected destinations, or destination credentials restart notification evaluation for the changed configuration. Disabling a monitor or all its destinations cancels queued messages and resets streaks. Re-enabling starts from fresh checks. Changing only the monitor or destination name preserves its incident state.

Delivery and credentials #

Streaks and incident state live in monitor_notification_state. Decisions and per-destination messages are committed together into notification_deliveries before network calls. A lease prevents concurrent schedulers from sending the same queued delivery simultaneously. Failed requests retry with exponential backoff and provider rate-limit delays, up to eight attempts. Permanent rejections stop retries. Each request has a 10-second timeout; delivery batches contain at most ten concurrent sends. Scheduler logs include delivery IDs and sanitized failures.

Delivery is at least once, not exactly once: a crash after a provider accepts a message but before the database records success can produce a duplicate. Disabling or editing a destination cancels queued messages but cannot retract a request already in flight.

API keys, tokens, passwords, and webhook URLs are write-only through the admin API. Editing with a blank secret retains the saved value; readable fields such as hosts, URLs, ports, recipients, and subjects remain visible for editing. Credentials are stored in PostgreSQL, so restrict database and backup access and use TLS for the admin panel. Public monitor and status-page responses do not expose notification configuration. Test messages use the same provider classes as scheduler messages.

Add another provider #

  1. Extend the provider/config schemas in packages/contracts/src/notifications.ts and add a migration for the allowed provider value.
  2. Add a separate provider module under packages/notifications/src/providers/, following telegram.ts and discord.ts. Its concrete class must extend NotificationProvider from ../notification-provider.ts, validate configuration in the constructor, and implement send(message: NotificationMessage): Promise<void>.
  3. Reuse message.ts for the shared notification text and delivery.ts for bounded HTTP requests and sanitized delivery errors. Keep provider-specific request bodies and response parsing in the provider module.
  4. Import the class into packages/notifications/src/index.ts, add it to the provider registry used by createNotificationProvider, and re-export it when direct construction is useful in tests. Add the provider's setup fields and redacted summary to the admin API/UI.
  5. Throw NotificationDeliveryError for delivery failures, with retryable and an optional retry delay. Never include credentials, URLs containing secrets, or raw provider responses in errors.
  6. Add provider tests using an injected fetch. The scheduler's incident and retry logic requires no provider-specific changes.

The scheduler integration tests use an isolated PGlite PostgreSQL instance and apply the repository migrations. They do not contact configured notification destinations or the deployment database.