Skip to main content
An alert rule watches one project, or one of its apps, and sends a message when something changes: when it starts firing and when it resolves, once each. Rules are checked every minute. Owners and admins manage alerts in Settings > Alerts; members see the rules and their history.

Rules

Nothing is decided below the minimum volume, so a quiet hour with two sessions and one crash is not a 50% crash rate; a rule that is firing stays firing until enough data shows it resolved. Without release, every release counts; without an endpoint, the worst endpoint with enough requests does. After a rule fires, it does not send again for its cooldown (60 minutes by default, 5 minutes to 7 days); a firing within it is recorded in the history but not sent. A rule can be muted for up to 30 days, and the same applies while it is. Changing a rule’s parameters or destination, or disabling it, starts it over without sending anything.

Destinations

  • Slack: an incoming webhook URL (in Slack, an app with Incoming Webhooks), https://hooks.slack.com/services/.... Messages name the project and release and link to the console.
  • Webhook: any https URL that answers 2xx. Each request is a signed JSON POST.
A destination’s URL is a secret (a Slack webhook URL can post to your channel), so Apsio stores it encrypted and never shows it again: the console shows its host. Use Send test to check one; the answer is a delivery status and, on failure, a code.

What a URL may be

Apsio refuses URLs that could reach its own network: the URL must be https, without a user or password, on port 443 or 1024 and up, and its host must resolve only to public addresses (no private, loopback, link-local, carrier-grade NAT, multicast or unique local address, and no IPv4-mapped IPv6). This is checked when the URL is saved and again before every delivery, and the request goes to the address that was checked. Redirects are not followed, and a destination has five seconds to answer.

Delivery codes

A timeout, connect, dns_refused or status_5xx is tried again on the next minutes, with the same event id, up to three attempts within an hour of the alert. Any other code is final.

Slow or unreachable destinations

A destination that answers slowly or not at all is paused, so it cannot hold back your other alerts:
  • Three answers in a row that are a timeout, connect, dns_refused or status_5xx, or that take longer than two seconds whatever they say, pause the destination for 5 minutes. Each further failure doubles the pause, up to 6 hours.
  • While it is paused, its alerts wait in the history as pending, without an attempt. One still waiting an hour after it was raised is recorded as failed, with its last code or expired.
  • When the pause runs out, the next alert is sent. A quick, successful delivery clears the pause and the count.
  • A successful Send test answered within two seconds also lifts a pause, unless the destination was paused for answering slowly: an endpoint could answer tests quickly and alerts slowly, so only a quick alert delivery lifts that one.
The console shows a paused destination with its last code. In the API, a destination has paused_until, consecutive_failures and last_error_code, which can also be slow.

History

The history lists everything alerting did, newest first: what fired and resolved, new issues, test sends, and what was held back. Each entry has its rule, destination, kind (fired, resolved, new_issue, test or suppressed), the measured value and detail, its delivery, its error_code, its attempts, and when it was raised and delivered. Read it in Settings > Alerts, or with GET /v1/orgs/{orgId}/alerts/events, filtered by project_id and paged with limit and before (the next_cursor of the previous page). Details carry numbers, ids and release names only.

Limits

An organization makes at most 30 delivery attempts a minute and 500 a day, retries included. Past that, a new alert is recorded in the history as suppressed (rate_limited), and each destination gets one message saying so per minute or day; a retry waits for a later minute. At most two deliveries of one organization are in flight at once. Test sends are limited to 5 a minute per destination.

Webhook requests

type is alert.fired, alert.resolved, alert.new_issue, alert.test or alert.suppressed. value is a rate from 0 to 1 (crash-free sessions, failed requests), the ratio to the hourly mean (spikes) or the occurrences (new issues). Messages carry names, ids and numbers only: never what your app’s users sent, such as error messages, stacks or request URLs.

Verify the signature

Each request is a POST with Content-Type: application/json, User-Agent: Apsio-Alerts/1 and two headers of its own:
t is when the request was signed, in Unix seconds. v1 is the hex HMAC-SHA256 of <t>.<raw body>, keyed with the destination’s signing secret: the whole whsec_... string, shown once when the destination is created or its secret rotated. To accept a request:
  1. Recompute the HMAC over the raw body, before parsing it, and compare it in constant time with each v1 in the header (today there is one).
  2. Reject a t more than five minutes from your clock, so an old request cannot be replayed.
  3. Keep the event ids you have handled and ignore a repeat: a retried delivery has the same id and a new t.
Answer 2xx within two seconds and do slow work after answering: a slower answer counts toward a pause, and after five seconds the request times out. In Node, pass the body as received (a string or a Buffer, such as express.raw() gives):
In Python, pass the body as bytes (in Flask, request.get_data()):
Both examples are run in CI against the signer Apsio uses. Rotating the secret (POST /v1/orgs/{orgId}/alerts/destinations/{destinationId}/rotate-secret) stops the old one at once.

API

Alerts are managed with the Console API under /v1/orgs/{orgId}/alerts: destinations, rules, test sends and the history.