AWS Backup & Disaster Recovery
Backups covering everything that matters, and one of them restored and timed in front of you.
From $449 2–5 days
Alerts that fire when customers are affected, and stay quiet the rest of the time.
From $299
Typically $299–$649, fixed in writing before anything starts.
What moves it up
Some of this you can check yourself, right now, for free: Uptime Monitoring Trial →
The problem is almost never that there is no monitoring. It is that the alarms which exist fire constantly and describe nothing: CPU above eighty per cent every weekday at nine, an alert that has been going off for six months, a channel everyone has muted. Meanwhile nothing at all watches the things that actually take a site down — a disk quietly filling, a certificate expiring on a Sunday, a queue that stopped being consumed an hour ago. The account has plenty of monitoring and no working alert, and those are different things.
aws cloudwatch describe-alarms --state-value ALARM — the ones that have been red for monthsaws cloudwatch describe-alarms --query 'MetricAlarms[?ActionsEnabled==`false`]' — the muted onesThe last real incident, and whether any alarm fired before a human noticeddf -h and certificate expiry on every endpoint, neither of which is usually alarmedProbably the thresholds and definitely what they are attached to. An alarm on CPU tells you a machine is busy, which is what machines are for; an alarm on error rate or p95 latency tells you a customer is having a bad time. The second is worth waking up for and the first is not.
Wherever you already look — email, Slack, SMS, PagerDuty. What matters more than the destination is that unacknowledged alerts escalate rather than sitting in a channel, because a notification nobody is responsible for is a notification nobody reads.
CloudWatch charges per metric, per alarm and per log gigabyte, and it is log ingestion that surprises people rather than the alarms. Retention gets set deliberately as part of this, because indefinite retention on a chatty application is the one line here that can genuinely get expensive.
Close to zero. If a set of alarms fires routinely and everything is fine, the alarms are wrong — and the real damage is not the noise, it is that the one that matters arrives in the same channel and gets skimmed past with the rest.
Backups covering everything that matters, and one of them restored and timed in front of you.
From $449 2–5 days
More than one of everything that matters, and a failover somebody has actually performed.
From $849 5–10 days
Tell me what you are running and I will come back with a fixed price and a date. If it turns out you do not need this, I will say that instead.
Prefer to talk? Book a free call ↗ · Or hire me on Upwork ↗ · Typical reply within one business day.
Sunday to Thursday, 09:00–18:00 EET. Outside that I will still look, but I will not promise a time.
One person, one time zone. If round-the-clock cover is what you need, you need a team, and I will say so rather than sell you a plan that cannot deliver it.
You pay Amazon directly and you keep control of the account. Nothing here resells your infrastructure or sits between you and your own billing.
Every service page lists exactly what pushes a quote above it, before you ask. You get a fixed number in writing before any work begins.