Updated: October 2026
A user swears they never signed in from another country, and the only witness is a log that rotated out last Tuesday. Logs are the one record of what happened on your machines, and they usually sit scattered across dozens of devices with default settings nobody chose. This guide covers log management end to end: what to collect, how to get it off the device, how long to keep it, and how to use it without drowning in it.
TL;DR
- Log management is the full lifecycle: generate, collect, parse, store, search and dispose. Having logs on each machine isn't the same thing.
- Start with a short list: Windows security events, PowerShell script block logging, Sysmon on servers, firewall and VPN syslog, and your identity provider's sign-in and audit logs.
- Get logs off the device fast. Windows Event Forwarding is built in, syslog covers network gear, and cloud logs need an API or export because default retention is short.
- Set retention from your regulator, insurer or contract first. Without one, a hot tier you can search in seconds plus a cheaper archive covers both troubleshooting and investigations.
- Log management stores and searches. A SIEM correlates and alerts on top of it. You need the first before the second earns its price.
What Is Log Management?
NIST's guide to computer security log management, SP 800-92, defines it as "the process for generating, transmitting, storing, analyzing, and disposing of computer security log data." That definition is from 2006. The problem it describes has only grown since.
Every device already writes logs. A Windows laptop writes to its event logs, a firewall writes to its own buffer, and Microsoft 365 keeps an audit trail in the cloud. None of that is log management. Log management starts when you decide which of those records matter, move them somewhere central, keep them long enough, and can search them when a ticket or an incident needs an answer.
The test is simple. Pick a question like "who added this account to Domain Admins, and from which machine?" If you can answer it in five minutes from one search box, you have log management. If the answer depends on remoting into three servers and hoping the Security log hasn't wrapped, you have logs.
NIST published a draft revision, SP 800-92 Rev. 1, the Cybersecurity Log Management Planning Guide, in October 2023. It's still an initial public draft, but it frames the work the same way: plan what you log, then build the infrastructure to support it.
CISA's introduction to log management is a good primer before you plan anything:
Where the Logs Come From
Five sources cover the questions a small IT team gets asked. Each answers a different kind of question, so skipping one leaves a blind spot rather than a thinner picture.
| Source | What it answers | Where it lives by default |
|---|---|---|
| Windows Security log | Who signed in, who got admin rights, which accounts and groups changed | Local event log on each machine, overwritten when full |
| PowerShell and Sysmon | What ran, with which command line, and what it connected to | Local operational event logs |
| Network gear (firewall, VPN, switches) | What left the network, who connected remotely, what got blocked | Device buffer, lost on reboot unless sent by syslog |
| Identity and cloud (Entra ID, Microsoft 365, Google Workspace) | Sign-ins, MFA changes, mailbox rules, file sharing | Provider portal, kept for a fixed number of days |
| Applications and servers | Errors, failed jobs, web requests, database access | Text files or app-specific logs |
Windows Event IDs Worth Collecting First
Windows logs a lot of events and very few of them matter on day one. Microsoft's own list of events to monitor ranks each by criticality, and the ones below cover the common investigation questions.
| Event ID | What it means | Why you want it |
|---|---|---|
| 4624 / 4625 | Successful / failed sign-in | Who got in, and who kept trying |
| 4740 | Account locked out | Password spraying and brute force show up here first |
| 4672 | Special privileges assigned to a new logon | An admin-level session started |
| 4720 | User account created | New accounts nobody requested |
| 4728 / 4732 | Member added to a global / local security group | Someone got added to an admin group |
| 4688 | New process created | What ran, if you enable the command line (see below) |
| 7045 / 4697 | Service installed | A common way to keep malware running |
| 4698 | Scheduled task created | Another common persistence method |
| 4719 | System audit policy changed | Someone turned your logging down |
| 1102 | Audit log cleared | Someone may be covering their tracks |
| 4104 | PowerShell script block logged | The script content, not just the fact it ran |
Two of these need a setting before they tell you anything useful. Event 4688 only carries the full command line when you enable "Include command line in process creation events" alongside the Audit Process Creation policy. Microsoft's own documentation warns that command lines can contain passwords, so restrict who can read the Security log once it's on. Event 4104 needs PowerShell script block logging turned on by Group Policy, and Microsoft recommends Protected Event Logging when you use it beyond diagnostics.
Sysmon adds depth where the Security log is thin. Microsoft's free Sysinternals tool logs process creation with full command lines and file hashes, driver loads, file creation, DNS queries and, if you configure it, network connections. Network connection logging is off by default, and so is image loading, because both are noisy. Sysmon doesn't analyze anything; it writes events to its own operational log for you to collect.
Getting Logs Off the Device
A log that only exists on the machine that wrote it disappears in three ways: the log fills up and overwrites itself, the machine gets reimaged, or an attacker clears it. Collection means getting a copy somewhere else, quickly.
Windows Event Forwarding
Windows Event Forwarding (WEF) is built into Windows, so there's no agent to deploy. Clients forward selected events to a Windows Event Collector (WEC) server. In a domain the connection is encrypted and mutually authenticated with Kerberos by default, even when you pick the HTTP option.
Microsoft recommends a push (source-initiated) subscription for scale. You point clients at the collector with a Group Policy, start the WinRM service, and add the Network Service account to the Event Log Readers group so the forwarder can read the Security log. Microsoft's guidance splits collection into two subscriptions: a baseline one for every device, and a targeted one with more events for machines you're investigating.
Plan for two limits. Microsoft's rule of thumb for a stable WEC server on commodity hardware is about 3,000 events per second across all subscriptions. And when a client's local log wraps before it reaches the collector, there's no notification and no marker in the stream. The local event log is WEF's only buffer, so size the Security log generously on laptops that spend days off the network.
One cheap win: WEF forwards "Rendered Text" by default, which includes the full event description and roughly doubles or triples each event's size. Switching a subscription to the "Events" format sends the compact XML instead, and Microsoft says it can more than double what one collector handles.
This walkthrough sets up a source-initiated subscription step by step:
Syslog for Network Gear and Linux
Firewalls, switches, VPN concentrators and Linux servers speak syslog. The current format is RFC 5424 from March 2009, which replaced the older BSD-style messages. Classic syslog runs over UDP port 514, and UDP drops messages silently when a collector is busy or a link is congested. For anything you'd need in an investigation, send syslog over TLS on TCP port 6514 (RFC 5425) if the device supports it.
Agents and Cloud APIs
Log shippers like the Elastic Agent, Grafana Alloy, Fluent Bit or a vendor's own agent tail files and event logs and send them on. They're the only practical option for application text logs, and they handle buffering better than WEF on machines that roam.
Cloud logs don't come to you by themselves. Entra ID and Microsoft 365 keep their logs in the provider's portal for a fixed period, then drop them. To keep them longer you export them: Entra diagnostic settings can send sign-in and audit logs to a storage account or a Log Analytics workspace, and Microsoft 365 audit data is reachable through the Management Activity API. Google Workspace exposes its audit data through the Admin SDK Reports API.
Parsing, Timestamps and Normalizing
Raw logs don't agree on anything. One device writes the user as "DOMAIN\jsmith", another as "jsmith@company.com", and a third only records an IP address. Parsing pulls the fields out of each format. Normalizing maps them to the same names, so one search for a user or a host finds every source.
Time is the field that breaks investigations. Sysmon writes its timestamps in UTC; many network devices log in local time; some don't sync their clocks at all. Keep everything on NTP, store timestamps in UTC, and convert only when you display them. A five-minute clock drift on a firewall is enough to make an attack chain look like it happened in the wrong order.
Common schemas like Elastic Common Schema (ECS) and the Open Cybersecurity Schema Framework (OCSF) exist so you don't invent field names from scratch. You don't need either on day one. You do need to pick one naming scheme and stick to it.
Enrichment is the last step. Adding the asset owner, the device role and a country for public IPs at ingest time turns a raw event into something a tired tech can read at 2 a.m.
How Long to Keep Logs
Retention is where good intentions quietly fail. The default settings keep far less than people assume, and upgrading later doesn't bring back what already expired.
What the Defaults Keep
Microsoft Entra ID keeps audit and sign-in logs for seven days on the free tier and 30 days on P1 or P2, according to Microsoft's data retention page. Microsoft is explicit that upgrading isn't retroactive: if you move from Free to P1, you only see the last seven days, and anything older is gone unless you archived it.
Microsoft 365 audit logs work differently. Audit (Standard) keeps records for 180 days; the default rose from 90 days for logs generated on or after 17 October 2023. Audit (Premium) keeps Exchange, SharePoint, OneDrive and Entra records for one year for users with an E5 license, and a separate add-on extends that to 10 years.
Windows event logs keep whatever fits in the file, then overwrite the oldest entries. On a busy server with default sizes, that can be days.
Setting Your Own Retention
Start with whatever you're bound by. Payment card, healthcare and government rules, cyber insurance questionnaires and client contracts each set their own minimum, and the strictest one wins. Write down which one applies and why, so the next person doesn't inherit a number nobody chose.
That inherited number is a common failure. In this r/sysadmin thread, an admin found an investigation reached back further than their retention, and nobody had picked the number on purpose:
The top reply suggests 90 days as a minimum and one year as their usual practice when no rule applies. Another commenter puts the useful window at about three months. The thread settles on a split: recent months for troubleshooting, older data for the rare investigation after a new indicator of compromise comes out.
That split is the hot and cold model. Keep a hot tier you can search in seconds for the window you troubleshoot in. Move older data to cheap storage you can restore in hours. Delete on schedule when the retention period ends, because logs you keep longer than required are one more thing to protect and disclose.
Protecting the Logs Themselves
A log an attacker can edit proves nothing. Send logs to a system the source machine can't write to after the fact, restrict who can delete them, and use immutable or write-once storage for the archive. Alert on event 1102 and on the Windows Event Log service stopping, since both are what clearing tracks looks like.
Searching and Alerting
Collecting logs is half the job. The other half is being able to ask a question and get an answer before the ticket goes stale.
A useful test for any setup is a handful of questions you should answer in minutes. Who signed in to this server in the last 24 hours? Which accounts were added to an admin group this month? Which machines ran this file hash? Where did this user sign in from before the password reset? If any of those needs a remote session to each machine, the collection or the parsing isn't done yet.
Alerting needs the opposite discipline. Alert on the few events that should almost never happen: the audit log cleared (1102), audit policy changed (4719), a new member in an admin group (4728 or 4732), a new service on a server that shouldn't get one (7045), or a burst of failed sign-ins and lockouts (4625 and 4740). Everything else stays searchable, not paged. Our guide to alert fatigue covers how to keep that list short enough that people still read it.
OpenFrame can run a script across a client's devices that checks audit policy, event log sizes and Sysmon status, and collect the output in one place, which is a fast way to find the machines that aren't logging what you think they are.
Log Management vs SIEM
People use the two terms as if they were the same product. They aren't, and the difference decides what you should buy first.
| Log management | SIEM | |
|---|---|---|
| Main job | Collect, store and search logs | Correlate events and raise security alerts |
| Question it answers | "What happened?" | "Is this an attack, right now?" |
| Typical users | IT and ops, plus security during investigations | Security analysts, a SOC or an MDR provider |
| Retention focus | Long, cheap, complete | Shorter hot window for fast correlation |
| Content you maintain | Parsers, dashboards, saved searches | Detection rules, use cases, tuning |
| Works without the other? | Yes | No, it depends on the logs arriving |
A SIEM sits on top of log management. It needs the same collection, parsing and time sync, then adds correlation rules and a response workflow. Buying a SIEM before your logs are complete and normalized means paying for correlation on incomplete data. Our guide to SIEM for MSPs covers when that step makes sense and what it costs to run.
Keeping the Volume and the Bill Under Control
Log platforms bill by what you send, what you store, or both. Volume creeps up with every source you add, so control it at the source instead of paying to store noise.
Filter where the log is written. A Sysmon configuration that excludes known-good processes, a WEF query that skips expected events, and a firewall set to log denies and admin changes instead of every allowed packet each cut volume before it costs anything. Microsoft's WEF baseline, for example, filters out access to the IPC$ and NetLogon shares because they're expected and noisy.
Send compact formats. The WEF "Events" format mentioned earlier is one example. Dropping duplicate fields and verbose debug output from application logs is another.
Tier by age. Keep the recent window hot and searchable, and move the rest to object storage where it costs a fraction as much. Restoring a month of cold logs for a rare investigation is cheaper than keeping a year hot.
Watch for the silent failures too. A collector that stops receiving from one site looks exactly like a quiet week. Alert when a source that normally sends logs goes silent for longer than it should.
Open-Source and Commercial Log Management Tools
You don't need a large budget to start. Several open-source stacks handle collection, storage and search for a small team, as long as someone owns them.
The ELK stack (Elasticsearch, Logstash and Kibana) is the long-standing option, and our ELK stack guide on OpenMSP walks through running it self-hosted. OpenSearch is the Apache-licensed fork, and Graylog builds a log management layer on top of OpenSearch or Elasticsearch with inputs, parsing and alerting.
Grafana Loki indexes labels instead of full text, which keeps storage cheap, and pairs with Grafana for dashboards. Our Grafana Loki review covers where that trade-off helps and where it hurts. Wazuh bundles agents, log analysis and security rules, which makes it closer to a SIEM than a plain log store.
On the commercial side, Splunk, Datadog, Sumo Logic and Microsoft Sentinel are common choices, each with its own pricing model based on ingest, retention or both. Price them against your actual daily volume after source filtering, not a vendor's estimate.
The open-source catch is setup and upkeep. In this r/sysadmin thread, an admin asking how to centralize Windows and Linux application logs got Loki, Wazuh, Graylog and OpenSearch suggestions, plus a warning that Graylog took one commenter about two years to master:
A Starter Plan for a Small IT Team
If you're starting from nothing, do it in this order. Each step is useful on its own, so you get value before the whole thing is finished.
| Step | What to do | Done when |
|---|---|---|
| 1. Decide retention | Write down which regulation, insurer or contract sets your minimum, and pick hot and cold periods | The number has an owner and a reason |
| 2. Export cloud identity logs | Send Entra ID (or Google Workspace) sign-in and audit logs to storage you control | Logs older than the default window exist somewhere |
| 3. Turn on the right Windows auditing | Apply an advanced audit policy, enable command line in 4688 and PowerShell script block logging | Events 4688 and 4104 show useful detail on a test machine |
| 4. Size local event logs | Raise the Security log size, especially on laptops | A week offline doesn't overwrite unforwarded events |
| 5. Stand up collection | WEF to a collector, or an agent, for Windows; syslog over TLS for network gear | Every server and firewall shows up as a source |
| 6. Add Sysmon to servers | Deploy with a tuned configuration, not the default | Process and DNS events arrive with hashes |
| 7. Normalize and sync time | NTP everywhere, UTC storage, consistent user and host fields | One search finds a user across every source |
| 8. Add five alerts | Log cleared, audit policy changed, admin group change, new service on servers, failed sign-in bursts | Each alert has a named owner |
| 9. Watch for silence | Alert when a normal source stops sending | You find out before an incident does |
| 10. Test it | Answer the four questions from the searching section in a timed drill | Each answer takes minutes, not hours |
Logs are also the first thing an incident responder asks for. Our incident response plan for small IT teams covers what happens next, and why the evidence you collect here decides how that goes.
The Short Version
Log management is the whole lifecycle, from choosing which events to write to deleting them on schedule. Collect a short, deliberate list first: Windows security events with command lines, PowerShell script blocks, Sysmon on servers, network syslog and cloud identity logs. Get them off the device quickly, keep them in UTC, and set retention from the rule you're bound by rather than whatever the default was. Once you can answer basic questions in minutes, a SIEM becomes worth considering, and our SIEM guide is the next read.

Aliaska Varieva
Head of Platform
Hi! I’m Aliaska, and I’ve been working as a software engineer (mostly Java + a bit Kotlin) for over 8 years now. I mostly spend my time building backend services, integrating systems, fixing bugs (the fun part 🙃), and making sure things don’t fall apart behind the scenes.
