About Expertise Work Managed Apps
Business Website Online Store Sales CRM Team Drive Online Academy Newsletter System Booking System Shared Inbox Knowledge Base Short Links Business Manager Photo Gallery Survey Platform Community Forum Project Boards Estate Agency Car Workshop Restaurant Clinic Photography Studio
AWS
Assess & advise Build & migrate Automate & operate Secure & comply Urgent & go-live
Projects
Hosted Monitoring & Dashboards Self-Hosted Observability Stack Bulk Document Data Extraction Email Deliverability Diagnosis & Repair SEO Migration Recovery AWS Security Review VPS Hardening & ModSecurity Cloud Architecture & Resilience Review SSL & Server Configuration Container Security Review DNS & Email Troubleshooting DevOps Deployment & Rollback Review WordPress Hardening Retainer Data Pipeline Rerun Review Metric Reconciliation
Free Tools
Website Health Check Email Domain Health Check DNS Health Check SSL Certificate Checker Redirect Chain Checker Robots.txt Checker XML Sitemap Validator Docker Compose Checker WordPress Security Check AWS IAM / S3 Policy Checker Domain Registration Lookup Uptime Monitoring Trial Downtime Cost Calculator AWS Cost Estimator Cloud Architecture Self-Assessment DevOps Engagement Builder Self-Managed VPS vs Managed AWS
Blog Certifications Hire Me

AWS Monitoring & Alerting

Alerts that fire when customers are affected, and stay quiet the rest of the time.

Price and scope

From $299

Typically $299–$649, fixed in writing before anything starts.

2–4 days
Working days, counted from the moment I have access — not from the day you agree.

What moves it up

  • Custom application metrics, rather than only what AWS emits on its own
  • On-call routing and escalation, where an unacknowledged alert has to go somewhere else
  • More than one environment, where staging noise must not reach a phone at night

Some of this you can check yourself, right now, for free: Uptime Monitoring Trial →

Alarms that fire, beside outages nobody is watchingOn the left, three alarms that fire constantly and have been muted. On the right, three conditions that actually cause outages — a filling disk, an expiring certificate and a queue that stopped being consumed — none of which has an alarm.FIRING, AND MUTEDCPU > 80%every weekday at 09:00disk > 70%red for six monthsstaging errorschannel is mutedWHAT ACTUALLY TAKES YOU DOWNdisk fullno alarmcertificate expiredno alarmqueue not consumedno alarmno overlap
AWS Monitoring & Alerting

What actually goes wrong

The problem is almost never that there is no monitoring. It is that the alarms which exist fire constantly and describe nothing: CPU above eighty per cent every weekday at nine, an alert that has been going off for six months, a channel everyone has muted. Meanwhile nothing at all watches the things that actually take a site down — a disk quietly filling, a certificate expiring on a Sunday, a queue that stopped being consumed an hour ago. The account has plenty of monitoring and no working alert, and those are different things.

How I find it

  • aws cloudwatch describe-alarms --state-value ALARM — the ones that have been red for months
  • aws cloudwatch describe-alarms --query 'MetricAlarms[?ActionsEnabled==`false`]' — the muted ones
  • The last real incident, and whether any alarm fired before a human noticed
  • df -h and certificate expiry on every endpoint, neither of which is usually alarmed

What you get

  • Alarms on symptoms a customer would feel — errors, latency, availability — not on raw resource metrics
  • Disk, certificate expiry and queue depth covered, because those are what actually cause outages
  • Every alarm routed somewhere a human is, with escalation if nobody acknowledges
  • Existing noisy alarms fixed or deleted, so the ones that remain mean something
  • A dashboard that answers "is it working" in one screen, not forty graphs

Questions

I already have CloudWatch alarms. What is different?

Probably the thresholds and definitely what they are attached to. An alarm on CPU tells you a machine is busy, which is what machines are for; an alarm on error rate or p95 latency tells you a customer is having a bad time. The second is worth waking up for and the first is not.

Where do the alerts actually go?

Wherever you already look — email, Slack, SMS, PagerDuty. What matters more than the destination is that unacknowledged alerts escalate rather than sitting in a channel, because a notification nobody is responsible for is a notification nobody reads.

Will this cost much to run?

CloudWatch charges per metric, per alarm and per log gigabyte, and it is log ingestion that surprises people rather than the alarms. Retention gets set deliberately as part of this, because indefinite retention on a chatty application is the one line here that can genuinely get expensive.

How many alerts should I expect in a normal week?

Close to zero. If a set of alarms fires routinely and everything is fine, the alarms are wrong — and the real damage is not the noise, it is that the one that matters arrives in the same channel and gets skimmed past with the rest.

Want this done?

Tell me what you are running and I will come back with a fixed price and a date. If it turns out you do not need this, I will say that instead.

Prefer to talk? Book a free call ↗  ·  Or hire me on Upwork ↗  ·  Typical reply within one business day.

When I answer

Sunday to Thursday, 09:00–18:00 EET. Outside that I will still look, but I will not promise a time.

No 24/7 desk, and I will not pretend otherwise

One person, one time zone. If round-the-clock cover is what you need, you need a team, and I will say so rather than sell you a plan that cannot deliver it.

Your AWS bill stays yours

You pay Amazon directly and you keep control of the account. Nothing here resells your infrastructure or sits between you and your own billing.

A price that starts with "from" is a starting price

Every service page lists exactly what pushes a quote above it, before you ask. You get a fixed number in writing before any work begins.