Operation
04.03 Set up once, then continuous

Monitoring & Alerts

You should not learn from your users that something is broken.

  • Errors
  • Uptime
  • Alerts

How it works

Monitoring is easy to set up and hard to tune. The hard part is deciding when it wakes somebody.

What we measure

Availability and response times of the interfaces. Errors in app and backend, with context rather than just a timestamp. The paths the business depends on — if nobody can place an order any more, that should show up even when everything is technically fine.

Alerts people take seriously

An alert that fires without cause is ignored within two weeks. At that point it is worse than none. We set a few clear thresholds and adjust them, rather than adding a new alert for every spike.

Urgent and important, kept apart

There are two channels. One wakes somebody, the other lands in a list for the next working day. You decide the assignment, not us — you know what costs you money.

What you can see yourself

You get access to the dashboard, not just a monthly report from us. How your software is doing should be visible without asking us.

What we need from you
  • A statement of what counts as an emergency and what can wait until tomorrow.
  • Who is reachable at night, if there is to be an on-call rota — and whether there should be one.
What you get

A dashboard with the numbers that matter to you, and alerts that fire only on real trouble.

If you skip this

Without monitoring, the first sign of an outage is an email from a customer. By then the problem has been running for hours and nobody knows since when.

Questions about this step?

A conversation costs nothing and takes half an hour. Afterwards you will know whether we are a fit.