Updated: October 2026
Your monitoring sends alerts all day, and a handful of them matter. The rest teach everyone to skim, and skimming is how the one that matters gets missed. Here's how to measure your alert load, cut noise at the source, and give what's left a clear owner.
What Is Alert Fatigue?
Alert fatigue is what happens when IT and security teams receive so many alerts, many of them low-value or false, that they start ignoring, delaying or auto-closing them. The volume trains people to treat every alert as noise, so the rare alert that signals a real outage or attack gets the same treatment.
The term comes from healthcare, where it's called alarm fatigue. The Joint Commission's 2013 sentinel event alert estimated that 85 to 99 percent of hospital alarm signals don't need clinical intervention. IT inherited the same pattern, just with disk space warnings instead of heart monitors.
What Ignored Alerts Cost
The obvious cost is the incident nobody saw. Splunk's State of Observability 2025, a vendor survey of 1,855 IT operations and engineering pros, found 73% had experienced outages caused by ignored or suppressed alerts.
Security teams see the same thing. Prophet Security's 2026 survey, also vendor-run, found 28% of alerts go uninvestigated. On a SIEM, that's where an attacker who knows how to stay at medium severity lives.
The quieter cost is people. A tech who spends a shift clearing the same CPU warning for the fifth client learns that alerts are chores. That lesson sticks long after the alerts get better.
This r/sysadmin thread is a good snapshot of where teams end up. RMM, SOC, backup, firewall and identity alerts all land in one place, and everyone tunes them out until one of the ignored ones bites. The top reply has a rule worth stealing: anyone who wants to add an alert has to be paged by it first.
Where the Noise Comes From
Alert fatigue in IT rarely comes from one tool. It comes from six or seven, each with sensible defaults, all firing into the same queue. Here's where the volume usually starts and the first thing to change.
| Alert source | Typical noise pattern | First fix |
|---|---|---|
| RMM performance (CPU, RAM, disk) | Point-in-time spikes that clear on their own | Alert only on sustained conditions |
| RMM availability | Every device behind a dead router alerts separately | Parent/child dependencies |
| Patching | Failed installs that succeed on the next retry | Alert after the second failure, not the first |
| Backup | Warnings on jobs that completed with skipped files | Split warnings from failures, route warnings to a daily digest |
| EDR | Low-severity detections on known admin tools | Exclusions scoped to specific paths and hashes |
| SIEM | Default rules nobody tuned for this environment | Raise or lower rule levels based on what the team acts on |
| Identity and M365 | Sign-in risk alerts for travelling users | Named locations and risk thresholds per group |
The pattern behind every row is the same. The default was written for a generic environment, and nobody revisited it after onboarding. That's a tuning job, and tuning jobs can be scheduled. If the monitoring side is new to you, our guide to remote monitoring and management covers how RMM alerting works.
Measure Your Alert Load Before You Tune
You can't tell whether tuning worked without a baseline. Pull 30 days of alert history from your RMM, PSA or SIEM, then run one formula:
Alerts per week × share that needed no action × minutes to triage each = time spent on noise
Here's an illustrative example. A team of four gets 2,400 alerts a week. That's 120 per tech per shift across a five-day week. If 70% needed no action and each takes 2 minutes to look at, the team spends 3,360 minutes a week on noise. That's 56 hours, or about 1.4 full-time techs, reading alerts that didn't matter.
The same export tells you where to start. Sort by alert type and count, and the noisiest types show themselves. Each one is a single tuning decision.
Keep the export. Run the same formula a month after tuning and compare. Track two more numbers alongside it: how many alerts each tech closes without doing anything, and how many incidents were first spotted by a user instead of an alert. The first should fall. The second should too.
Tuning Moves That Cut Volume
Google's SRE book puts the goal in one line: "Every page should be actionable." The moves below get you closer, and every monitoring tool worth running supports some version of them.
- Alert on duration, not a single reading. Datto RMM lets you set how long a condition must hold before it alerts, from 1 to 60 minutes. Prometheus uses a
for:clause for the same job. - Deduplicate and group. Collapse repeats of the same alert into one, with a count. Alertmanager groups related alerts and waits before sending, so a burst becomes one notification.
- Suppress children when the parent is down. Zabbix trigger dependencies and PRTG dependencies both stop 40 devices from alerting when one router is the problem.
- Use maintenance windows. Patch night shouldn't page anyone about reboots you scheduled.
- Auto-resolve on recovery. If the condition clears, close the alert. Datto RMM can resolve an alert automatically after a set time without a retrigger.
- Baseline per device, not globally. A build server at 90% CPU is normal. A receptionist's laptop at 90% isn't.
- Tune SIEM rule levels. Wazuh lets you override a built-in rule's level in your local rules, so a noisy detection drops below your alert threshold without being deleted.
- Script the known fixes. Disk cleanup and service restarts don't need a human to decide. They need a human to approve the script once.
This r/msp thread shows what these look like in a real stack. The replies trade concrete settings: alert on 95% CPU only after 5 minutes, close the ticket when the condition clears, and move devices that run hot for days onto an upgrade list instead of the ticket queue.
An Alert Triage Matrix
Tuning cuts volume. Ownership decides what happens to what's left. Every alert class needs a severity, an owner and a response window, written down before the alert fires.
| Severity | Example | Owner | Response window | Automated first step |
|---|---|---|---|---|
| Critical | Server or site offline, ransomware detection | On-call tech, paged | 15 minutes | Open incident, attach device context |
| High | Backup failed twice, EDR high-severity detection | Service desk queue, flagged | 1 hour | Retry job, isolate device if policy allows |
| Medium | Disk above 90% for 30 minutes | Service desk queue | Same business day | Run disk cleanup script |
| Low | Patch installed after retry, single sign-in risk | Daily digest | Next review | None, log only |
| Info | Agent check-in, job success | Dashboard only | None | None |
The bottom two rows are where the biggest wins hide. Anything that doesn't need a human today shouldn't reach a human's queue today. A daily digest keeps the record without spending anyone's attention.
Run the earlier example through both steps. Deduplication and grouping might take 2,400 weekly alerts down to 1,100. Of those, the 720 that need action go to the right queue, and around 90 reach the critical and high rows. That's the pile a person should see.
Where AI Triage Fits
Tuning has a ceiling. Anton Chuvakin at Google Cloud put it plainly: "alert fatigue is not one problem." Some of it is false positives, which tuning fixes. Some is missing context, or too few people for the volume, which it doesn't.
That's where AI earns its place. It can group related alerts into one incident, pull in the device, client and recent changes, and draft the first response. Your techs read 30 incidents instead of 300 alerts, and spend the time they get back on the ones that need judgement. The AI does the sorting. People still make the call.
Set the guardrails before you switch it on. Decide which alert classes AI may group or close, which fixes it may suggest, and which always need a human. Start with the low and info rows of the matrix, check its work for a few weeks, then widen the scope.
OpenFrame, Flamingo's open, AI-native infrastructure layer for IT and security, keeps the native RMM agent, a SIEM module and the script library in one place. Nothing risky runs without a technician's approval. Flamingo's published task pricing puts alert triage at 30 minutes and $80 of technician time by hand, against about 2 minutes and $0.06 with its Fae and Mingo agents. For the wider view on where this is heading, see our take on AI in MSP work.
This short explainer from Dr. K Cybersecurity covers the security side of the same problem, and makes a point worth repeating: missed alerts come from volume and priorities, not lazy analysts.
Start With One Alert Type
Alert fatigue is a tuning and ownership problem, and both can go on a calendar. Pull 30 days of history, find the noisiest alert type, and give it a duration condition, a dependency or a digest this week. Then do the next one.
If your alerts end up in a security queue, our guide to the SIEM alert pipeline is the next read.

"Fae" Grace Meadows
Lead AI Fairy
Some things defy easy explanation: magic dust, the northern lights… and Flamingo’s AI Angels. Think Charlie’s Angels, reimagined with automation brains and serious RMM (Remote Monitoring & Management) chops. Weird? A little. Effective? Absolutely. That’s the job.
