Flamingo Raises $4.5M Seed Round

Skip to content

Updated: October 2026

Nobody gets paged when a client's Azure bill doubles overnight or a certificate quietly expires in a key vault. The server closet has a UPS that beeps, and a cloud tenant has whatever someone remembered to switch on. Here's what cloud based monitoring covers, which native tools do the work in Azure, AWS and Google Cloud, and what to alert on first.

What Cloud Monitoring Covers That On-Prem Doesn't

On-prem monitoring watches boxes you own. CPU, disk, memory, the switch port, the UPS. If the hardware is healthy and the service answers, you're mostly done.

In the cloud you don't own the hardware, so the list changes. You still watch your own resources, but three new signals join them. The provider's own health is one: an outage in a region you run in is your outage too. Cost is another, because a runaway resource shows up on the invoice before it shows up anywhere else. Identity is the third, since in a tenant the sign-in log is the front door.

SaaS adds a fourth layer. Microsoft 365 or Google Workspace can be down or degraded while every VM you run is green. That still lands on your help desk.

For the network side of hybrid setups, network performance tools still do the job. Our guide to IT infrastructure monitoring covers the on-prem layers in depth.

The Native Cloud Monitoring Tools

Each big cloud ships its own monitoring stack. They're the first place to start because the data is already there.

CloudNative toolWhat it collectsProvider health
AzureAzure MonitorMetrics, logs, traces; logs land in a Log Analytics workspace (KQL)Azure Service Health
AWSAmazon CloudWatchMetrics, logs, alarms, dashboardsAWS Health
Google CloudCloud Monitoring (Google Cloud Observability)Metrics for most Google Cloud services, with Cloud Logging alongsidePersonalized Service Health
Microsoft 365Service health in the admin centerIncidents and advisories for your tenantSame page

Microsoft describes Azure Monitor as its "unified observability service" for cloud and hybrid environments. It collects logs and metrics from Azure resources, including Entra ID audit logs. Logs and traces go to Log Analytics workspaces, and Prometheus-style metrics go to a separate Azure Monitor workspace. Azure Service Health sits next to it and sends alerts about outages, planned maintenance and health advisories for the regions you use.

Amazon CloudWatch covers metrics, logs, alarms and dashboards for AWS resources. Many AWS services send basic metrics to it for free. Cross-account observability lets one central account watch metrics and logs from several source accounts.

Google Cloud Monitoring is part of Google Cloud Observability, next to Cloud Logging and Cloud Trace. It can also pull in metrics from AWS accounts, which helps if you run both.

For Microsoft 365, go to Health > Service health in the admin center. You can get email notifications there for up to two addresses per admin. Planned maintenance isn't shown on that page; it lives in the Message center.

What to Alert On First

A new tenant can generate hundreds of alerts on day one. Start with the few that prevent the calls you dread, then add resource alerts as you learn what normal looks like.

  1. Provider health for your regions. An Azure Service Health or AWS Health alert tells you it's not you before the tickets start.
  2. Budgets. AWS Budgets can alert on actual and forecasted spend. AWS notes a delay between usage and billing, so treat it as an early warning, not a hard stop.
  3. Quotas. Service Quotas supports CloudWatch alarms when usage nears a limit. Hitting a quota at 2am looks exactly like an outage.
  4. Expiring certificates and secrets. ACM sends daily events starting 45 days before expiry for private and imported certificates, and 30 days for public ones. Key Vault can emit a near-expiry event 30 days out through Event Grid, but only if you subscribe first.
  5. Identity. Alert on new admin role assignments and sign-ins that break your normal pattern. The sign-in log is the front door, so it needs a watcher.

Only then add the classic resource alerts: CPU, memory, disk, failed requests. Microsoft's Azure Monitor Baseline Alerts project gives a starting set. The r/AZURE thread below has a fair warning about it: set your thresholds before you deploy, or the inbox fills up.

Every alert should have an owner and an action. If nobody would act on it, it belongs on a dashboard. Our guide to alert fatigue covers how to trim the rest.

Keep the Logs Where You Can Find Them

Alerts tell you something happened. Logs tell you what. In the cloud, the default retention is often shorter than the time it takes someone to notice a problem.

Entra ID is the usual example. Sign-in and audit logs stay for 7 days on the free tier and 30 days on P1 or P2. Microsoft notes that upgrading doesn't bring back older data. If a compromised account surfaces in week three, a free tenant has nothing left to show you. The fix is to route the logs to a Log Analytics workspace or a storage account from day one.

Microsoft 365 adds a second trail. The Unified Audit Log is separate from the Entra logs, and its retention is managed in Microsoft Purview. Check both before you promise a client a 90-day look-back.

Retention costs money, so decide it per client and write it down. A short default with a longer archive for identity and admin activity covers the questions that come up after an incident.

Native Tools vs Third-Party Platforms

Native tools win on depth and price for one cloud and one tenant. The data is already collected, the alerts are close to the resource, and there's no extra agent.

Third-party platforms fall into a few groups:

  • Observability platforms like Datadog, Grafana Cloud or New Relic pull metrics, logs and traces from several clouds into one place.
  • RMM and infrastructure monitoring tools add cloud checks next to servers and endpoints, which suits teams that live in one console.
  • Cost and FinOps tools focus on spend, tagging and forecasts across accounts.

The tipping point is usually tenants, not clouds. In this r/sysadmin thread, a small MSP asks how to run Azure Monitor for clients at scale. One reply says they tried Azure-native tooling through Lighthouse and switched to a third-party platform once they passed about 20 customers. The logs still stay in each customer's own workspace. The third-party tool queries them.

That's a pattern worth copying either way. Keep each client's data in the client's tenant, and bring the alerts to you.

The Multi-Cloud Problem

Run Azure and AWS side by side and you run two monitoring languages. Azure logs speak KQL. CloudWatch Logs Insights speaks its own query language, SQL or PPL. Alert rules, action groups and notification channels all work differently.

OpenTelemetry takes some of the sting out. Both Azure Monitor and CloudWatch accept OpenTelemetry data, so you can instrument once and send the same telemetry to more than one place. Google Cloud Monitoring can read AWS metrics, and Azure Arc brings servers in other clouds into Azure Monitor.

None of that removes the on-call question. Someone still has to know which console to open at 2am, and the answer shouldn't depend on which cloud broke.

A Short Checklist for Choosing Cloud Monitoring Tools

  1. List every cloud, tenant and SaaS app you're responsible for, including the Microsoft 365 tenant.
  2. Turn on provider health alerts for the regions you use before anything else.
  3. Set budget and quota alerts on every subscription and account.
  4. Find every certificate and secret with an expiry date, and route those events to a real person.
  5. Decide where logs live and how long, especially identity logs.
  6. Pick native tools for one tenant; look at a third-party platform once tenants or clouds multiply.

Cloud monitoring adds provider health, cost and identity to the list you already watch. Start with those five alerts, keep the data in each tenant, and add resource alerts as your baseline settles. Next, read our guide to IT operations management tools for the wider tool picture.

Conrad Lunderstedt

Conrad Lunderstedt

Solution Architect

I'm Conrad, Solution Architect at Flamingo. I've spent about 26 years in IT, roughly half of it inside MSPs and the rest in enterprise environments, so I've watched vendor decisions get made on both sides of that line. Now I spend my days talking with MSPs about the stack they already run, and helping them work through the requests and issues that come with it.

Related Content

Blog Posts

Product Releases

Podcasts

Webinars

Case Studies

Events

Onboarding Guides

Frequently Asked Questions

Cloud Monitoring Tools

Cloud based monitoring tracks the health, performance and cost of resources running in a public cloud such as Azure, AWS or Google Cloud, plus SaaS services like Microsoft 365. Beyond CPU, memory and logs, it watches signals on-prem monitoring never had: the provider's own service health, spend against budget, quotas and identity sign-ins.
Each cloud ships a native tool: Azure Monitor with Log Analytics and Azure Service Health, Amazon CloudWatch with AWS Health, and Google Cloud Monitoring as part of Google Cloud Observability. Microsoft 365 service health lives in the admin center. Third-party observability, RMM and FinOps platforms add a single view across several clouds or tenants.
Start with provider service health for the regions you use, budget alerts on actual and forecasted spend, quota alarms, expiring certificates and secrets, and identity events such as new admin role assignments. Add CPU, memory and disk alerts once you know what normal looks like, and give every alert an owner and an action.
Not for one cloud and one tenant; native tools are deeper and cheaper there. Once you manage several tenants or several clouds, each with its own console and query language, a third-party platform that queries each tenant's own logs and raises alerts in one queue usually saves more time than it costs.

MSP AI Agents

On a five-person desk, reported deployments show $78,000 to $130,000 in annual direct labor savings, roughly 30% fewer escalations, and 15% to 20% better SLA compliance. Broader MSP adoption data adds ticket handling time cut by 45% and five to 12 points of margin, all from reclaimed capacity rather than headcount cuts.
Yes. In production MSP shops today, 10% to 25% of tickets close before a human opens them. Thread alone has processed 173 million tickets across 750-plus MSP partners at 96% triage accuracy, handing back 490,000-plus technician hours. Agents own the low-risk, high-volume work (password resets, MFA enrollment, known installs, onboarding and offboarding) and flag anything that touches production data or needs judgment for a human to take.

About OpenFrame

OpenFrame isn't built to plug into your stack. It replaces it. Instead of duct-taping a dozen tools together (RMM, MDM, SIEM, patching, remote access, each its own login and bill), we bundle it into one unified platform: RMM, MDM, monitoring, automation, remote access, patch management, security monitoring, and ticketing, plus built-in AI copilots. So "does it integrate with X?" usually means: you won't need X anymore.
Most platforms give you one piece and expect you to bolt the rest on. OpenFrame unifies the whole stack in one place, with AI copilots built in. Fewer logins, fewer bills, less duct tape.
In the cloud, on US soil. Your data stays stateside.
Both. It's built for MSPs and MSSPs alike.