Updated: October 2026
Nobody gets paged when a client's Azure bill doubles overnight or a certificate quietly expires in a key vault. The server closet has a UPS that beeps, and a cloud tenant has whatever someone remembered to switch on. Here's what cloud based monitoring covers, which native tools do the work in Azure, AWS and Google Cloud, and what to alert on first.
What Cloud Monitoring Covers That On-Prem Doesn't
On-prem monitoring watches boxes you own. CPU, disk, memory, the switch port, the UPS. If the hardware is healthy and the service answers, you're mostly done.
In the cloud you don't own the hardware, so the list changes. You still watch your own resources, but three new signals join them. The provider's own health is one: an outage in a region you run in is your outage too. Cost is another, because a runaway resource shows up on the invoice before it shows up anywhere else. Identity is the third, since in a tenant the sign-in log is the front door.
SaaS adds a fourth layer. Microsoft 365 or Google Workspace can be down or degraded while every VM you run is green. That still lands on your help desk.
For the network side of hybrid setups, network performance tools still do the job. Our guide to IT infrastructure monitoring covers the on-prem layers in depth.
The Native Cloud Monitoring Tools
Each big cloud ships its own monitoring stack. They're the first place to start because the data is already there.
| Cloud | Native tool | What it collects | Provider health |
|---|---|---|---|
| Azure | Azure Monitor | Metrics, logs, traces; logs land in a Log Analytics workspace (KQL) | Azure Service Health |
| AWS | Amazon CloudWatch | Metrics, logs, alarms, dashboards | AWS Health |
| Google Cloud | Cloud Monitoring (Google Cloud Observability) | Metrics for most Google Cloud services, with Cloud Logging alongside | Personalized Service Health |
| Microsoft 365 | Service health in the admin center | Incidents and advisories for your tenant | Same page |
Microsoft describes Azure Monitor as its "unified observability service" for cloud and hybrid environments. It collects logs and metrics from Azure resources, including Entra ID audit logs. Logs and traces go to Log Analytics workspaces, and Prometheus-style metrics go to a separate Azure Monitor workspace. Azure Service Health sits next to it and sends alerts about outages, planned maintenance and health advisories for the regions you use.
Amazon CloudWatch covers metrics, logs, alarms and dashboards for AWS resources. Many AWS services send basic metrics to it for free. Cross-account observability lets one central account watch metrics and logs from several source accounts.
Google Cloud Monitoring is part of Google Cloud Observability, next to Cloud Logging and Cloud Trace. It can also pull in metrics from AWS accounts, which helps if you run both.
For Microsoft 365, go to Health > Service health in the admin center. You can get email notifications there for up to two addresses per admin. Planned maintenance isn't shown on that page; it lives in the Message center.
What to Alert On First
A new tenant can generate hundreds of alerts on day one. Start with the few that prevent the calls you dread, then add resource alerts as you learn what normal looks like.
- Provider health for your regions. An Azure Service Health or AWS Health alert tells you it's not you before the tickets start.
- Budgets. AWS Budgets can alert on actual and forecasted spend. AWS notes a delay between usage and billing, so treat it as an early warning, not a hard stop.
- Quotas. Service Quotas supports CloudWatch alarms when usage nears a limit. Hitting a quota at 2am looks exactly like an outage.
- Expiring certificates and secrets. ACM sends daily events starting 45 days before expiry for private and imported certificates, and 30 days for public ones. Key Vault can emit a near-expiry event 30 days out through Event Grid, but only if you subscribe first.
- Identity. Alert on new admin role assignments and sign-ins that break your normal pattern. The sign-in log is the front door, so it needs a watcher.
Only then add the classic resource alerts: CPU, memory, disk, failed requests. Microsoft's Azure Monitor Baseline Alerts project gives a starting set. The r/AZURE thread below has a fair warning about it: set your thresholds before you deploy, or the inbox fills up.
Every alert should have an owner and an action. If nobody would act on it, it belongs on a dashboard. Our guide to alert fatigue covers how to trim the rest.
Keep the Logs Where You Can Find Them
Alerts tell you something happened. Logs tell you what. In the cloud, the default retention is often shorter than the time it takes someone to notice a problem.
Entra ID is the usual example. Sign-in and audit logs stay for 7 days on the free tier and 30 days on P1 or P2. Microsoft notes that upgrading doesn't bring back older data. If a compromised account surfaces in week three, a free tenant has nothing left to show you. The fix is to route the logs to a Log Analytics workspace or a storage account from day one.
Microsoft 365 adds a second trail. The Unified Audit Log is separate from the Entra logs, and its retention is managed in Microsoft Purview. Check both before you promise a client a 90-day look-back.
Retention costs money, so decide it per client and write it down. A short default with a longer archive for identity and admin activity covers the questions that come up after an incident.
Native Tools vs Third-Party Platforms
Native tools win on depth and price for one cloud and one tenant. The data is already collected, the alerts are close to the resource, and there's no extra agent.
Third-party platforms fall into a few groups:
- Observability platforms like Datadog, Grafana Cloud or New Relic pull metrics, logs and traces from several clouds into one place.
- RMM and infrastructure monitoring tools add cloud checks next to servers and endpoints, which suits teams that live in one console.
- Cost and FinOps tools focus on spend, tagging and forecasts across accounts.
The tipping point is usually tenants, not clouds. In this r/sysadmin thread, a small MSP asks how to run Azure Monitor for clients at scale. One reply says they tried Azure-native tooling through Lighthouse and switched to a third-party platform once they passed about 20 customers. The logs still stay in each customer's own workspace. The third-party tool queries them.
That's a pattern worth copying either way. Keep each client's data in the client's tenant, and bring the alerts to you.
The Multi-Cloud Problem
Run Azure and AWS side by side and you run two monitoring languages. Azure logs speak KQL. CloudWatch Logs Insights speaks its own query language, SQL or PPL. Alert rules, action groups and notification channels all work differently.
OpenTelemetry takes some of the sting out. Both Azure Monitor and CloudWatch accept OpenTelemetry data, so you can instrument once and send the same telemetry to more than one place. Google Cloud Monitoring can read AWS metrics, and Azure Arc brings servers in other clouds into Azure Monitor.
None of that removes the on-call question. Someone still has to know which console to open at 2am, and the answer shouldn't depend on which cloud broke.
A Short Checklist for Choosing Cloud Monitoring Tools
- List every cloud, tenant and SaaS app you're responsible for, including the Microsoft 365 tenant.
- Turn on provider health alerts for the regions you use before anything else.
- Set budget and quota alerts on every subscription and account.
- Find every certificate and secret with an expiry date, and route those events to a real person.
- Decide where logs live and how long, especially identity logs.
- Pick native tools for one tenant; look at a third-party platform once tenants or clouds multiply.
Cloud monitoring adds provider health, cost and identity to the list you already watch. Start with those five alerts, keep the data in each tenant, and add resource alerts as your baseline settles. Next, read our guide to IT operations management tools for the wider tool picture.
Conrad Lunderstedt
Solution Architect
I'm Conrad, Solution Architect at Flamingo. I've spent about 26 years in IT, roughly half of it inside MSPs and the rest in enterprise environments, so I've watched vendor decisions get made on both sides of that line. Now I spend my days talking with MSPs about the stack they already run, and helping them work through the requests and issues that come with it.
