# Alerts

An alert should be rare, true and clear. OpenPing confirms a failure before it tells anyone, groups failures with one cause into one incident, and writes the message in plain words.

## What an alert says

```text
Checkout is down
From Mumbai and Frankfurt since 14:05. It started 3 minutes after the 14:02 deploy.
```

What failed, where, since when, and what changed just before. Never a status code with no context.

## Where alerts go

| Channel | How it works |
| --- | --- |
| Email | On from the moment you sign in, to your own address. Add up to 20 addresses per channel. |
| Slack | Paste an incoming-webhook URL for the channel. The message has the state in words, since when, the regions, and a link to the incident. |
| Webhook | A signed JSON `POST` to your URL, retried if it fails. |
| Push to phone | Through the OpenPing phone app, once you have paired it. |

Press **Send a test alert** after adding a channel. It tells you how long each one took to arrive. Microsoft Teams, Discord, Telegram, SMS, phone calls and PagerDuty are planned.

## Alert rules

A rule says who hears about what, and after how long. It has steps: the first is told at once, each later step only if nobody has acknowledged the incident by then.

- **Steps:** up to 6, each with channels and “after N minutes”.
- **Severities:** by default `critical` and `normal` monitors alert; `low` ones go into a daily digest.
- **Reminders:** while an incident stays open and unacknowledged, a reminder every 60 minutes unless you change it.

A monitor uses its own rule if it has one, otherwise your organisation’s default rule. New accounts start with a default rule that emails the person who signed up.

## How it stays calm

- **Confirm first.** No alert until other regions have confirmed the failure: two of three must fail. See [how a failure is confirmed](https://openping.ai/docs/monitors#confirm).
- **One incident per cause.** Monitors on the same host join one incident. So does a monitor that depends on one already in an incident: if the gateway is down, the agent tests that use it fold in.
- **Our problem, not yours.** If many unrelated monitors fail from one region at once, that region is set aside for a while and nobody is paged.
- **No repeats.** A flapping monitor sends one alert. Each channel gets at most 30 messages a minute; the rest fold into one “and N more”.
- **Rate limits are not outages.** A 429 from a model provider is counted on its own and never fails a monitor by itself.

## Incidents

An incident opens by itself when a critical or normal monitor goes down, or by hand. Its timeline fills itself: first failure, confirmations by region, alerts sent, who acknowledged, nearby deploys, recovery. A short summary at the top says what is broken, where and since when. It never names a cause it can’t point to.

States match the status page: Investigating, Identified, Monitoring, Resolved.

## Webhooks

Each webhook carries an `OpenPing-Signature` header: a timestamp and an HMAC-SHA256 of the body, made with the signing key you were shown once when you created the channel.

```http
POST /your/endpoint
OpenPing-Signature: t=1791100000,v1=5f2b…

{
  "id": "evt_…",
  "type": "monitor.down",
  "at": "2026-10-04T14:05:00Z",
  "title": "Checkout is down",
  "body": "From Mumbai and Frankfurt since 14:05.",
  "severity": "critical",
  "monitor": { … },
  "incident": { … }
}
```

Event types: `monitor.down`, `monitor.degraded`, `monitor.up`, `incident.opened`, `incident.updated`, `incident.resolved`, `mcp.tools_changed`, `suite.failed`, `certificate.expiring` and `heartbeat.late`.
