Flamingo Raises $4.5M Seed Round

Skip to content

Updated: October 2026

A client says "we have an SLA with you," and nobody on the team can say what it promises. Half the arguments about slow tickets start there. Here's what an SLA means in business, how it differs from an OLA and an SLO, and how to write and measure one your team can keep.

What Is an SLA?

A service level agreement (SLA) is the part of a contract that says how well a service will be delivered, not just what the service is. It turns "we'll support your IT" into numbers both sides can check: how fast someone responds, how fast a fix lands, how much uptime counts as normal.

A usable SLA answers seven questions:

  • Scope. Which services, users and devices are covered, and which aren't.
  • Service hours. When the clock runs: business hours, extended hours or 24/7.
  • Priorities. How a ticket gets its P1 to P4 label.
  • Targets. Response and resolution times per priority, plus any uptime target.
  • Measurement. Which system records the times, and what pauses the clock.
  • Remedies. What happens when a target is missed: service credits, a review, an exit clause.
  • Review. How often both sides look at the numbers and change them.

If one of those is missing, the gap gets filled during an outage, by whoever is angriest.

SLA vs OLA vs Underpinning Contract

An SLA is the promise to the customer. Keeping it usually depends on people who never signed it.

An operational level agreement (OLA) is an internal agreement between teams. It "defines interdependent relationships in support of a service-level agreement," in Wikipedia's wording of the ITIL idea. If the service desk promises a four-hour fix for a P2, the network team needs an OLA that says it will pick up an escalation within one hour.

An underpinning contract (UC) does the same job with an outside supplier: the ISP, the hardware vendor, the SaaS platform. When the internet line is down, your resolution time is only as good as the carrier's repair commitment.

Write the three in order. Set the customer SLA, then check every internal team and supplier it depends on can hit its share. An SLA with no OLA or UC behind it is a guess.

SLA vs SLO vs SLI

These three sound alike and get used interchangeably. Google's Site Reliability Engineering book pins them down.

An SLI is "a carefully defined quantitative measure of some aspect of the level of service," like the share of successful logins or median ticket response time. An SLO is "a target value or range of values for a service level that is measured by an SLI," like 95% of P2 tickets answered within an hour. An SLA is "an explicit or implicit contract with your users that includes consequences of meeting (or missing) the SLOs they contain."

So the SLI is the measurement, the SLO is the goal, and the SLA is the promise with consequences attached. Teams often run internal SLOs a notch tighter than the SLA, so they get warned before the contract is breached.

Google Cloud's short explainer walks through the same split with examples:

Response Time vs Resolution Time

Response time is how long until a person acknowledges the ticket and starts work. Resolution time is how long until the problem is fixed and the user confirms it. They measure different things, and SLAs that blur them cause fights.

Response is fully in your control, so it's easy to promise. Resolution isn't. You can't guarantee a fix when you don't yet know the cause, or when the fix depends on the carrier. That's why a common pattern is to commit hard to response and treat resolution as a target with named exceptions.

This r/sysadmin thread on P1 and P2 targets lands on the same split. The top reply calls resolution "a hard thing to put on an SLA" and suggests a one-hour maximum to start work on a P1:

Two rules decide whether the numbers mean anything. First, say which clock runs: a four-hour target on business hours can stretch over a weekend, while a 24/7 clock can't. Second, list what stops the clock. "Waiting on customer" and "waiting on third party" usually pause it, and the SLA should say so in writing.

A Sample SLA Table

Priority comes from impact and urgency: how many people are affected, and how fast the damage grows. Our guide to the incident management process covers the priority matrix in full. Once priorities exist, the targets fit in one table:

PriorityExampleResponseResolution targetClock
P1 CriticalWhole office can't work, server or line down30 minutes4 hours24/7
P2 HighA team or a key app is down, workaround exists1 hour8 business hoursBusiness hours
P3 NormalOne user affected, can keep working4 business hours3 business daysBusiness hours
P4 LowRequests, how-to questions, new setups1 business day5 business daysBusiness hours

These numbers are a starting point, not a standard. The right target is what the client's downtime costs, weighed against what faster cover costs you. Tighter numbers need more people on call, and that belongs in the price.

Service Credits and Penalties

The consequence is what makes an SLA more than a wish. The most common form is a service credit: a discount on the next invoice when a target is missed.

Cloud providers publish theirs. The Amazon EC2 SLA (last updated May 2022) commits to 99.99% monthly uptime for instances spread across availability zones. Drop below that and the credit is 10% of the bill, rising to 30% below 99.0% and 100% below 95.0%. That 99.99% leaves room for about 4.3 minutes of downtime in a 30-day month.

For a service desk, credits usually attach to response times, because those are fully in your hands. Keep the remedy proportional. A credit that wipes out a month's margin for one missed P3 makes the team defensive, and defensive teams game metrics.

Measuring SLAs Without Gaming Them

Measure SLAs in the PSA or ticketing system, not in a spreadsheet someone fills in on Friday. The system stamps created, first response, status changes and resolved times, and the report comes straight from those stamps.

Then watch for the ways the numbers get bent. An auto-reply that counts as "first response." Tickets parked in "waiting on customer" with no question asked. P2s quietly downgraded to P3. Tickets closed and reopened to restart the clock. Each one keeps the dashboard green while users wait.

Pair the SLA numbers with signals that are harder to fake: reopen rate, customer satisfaction on closed tickets, and a monthly sample of breached tickets read by a human. This r/msp debate on "blind adherence" is a good reminder that an SLA is a floor. The top answer: when people are free, low-priority work gets done right away.

SLAs in Short

An SLA turns a service into numbers: response and resolution times per priority, the clock they run on, what pauses it, and what happens when a target is missed. OLAs and supplier contracts are what make those numbers keepable, and SLOs give the team an early warning. Measure it in the ticketing system and pair it with signals nobody can game.

If you're on the buying side, our list of questions to ask an MSP covers how to read an SLA before you sign.

Conrad Lunderstedt

Conrad Lunderstedt

Solution Architect

I'm Conrad, Solution Architect at Flamingo. I've spent about 26 years in IT, roughly half of it inside MSPs and the rest in enterprise environments, so I've watched vendor decisions get made on both sides of that line. Now I spend my days talking with MSPs about the stack they already run, and helping them work through the requests and issues that come with it.

Related Content

Blog Posts

Product Releases

Podcasts

Webinars

Case Studies

Events

Onboarding Guides

Frequently Asked Questions

Service Level Agreements

SLA stands for service level agreement. It is the part of a contract that sets how well a service is delivered, in numbers both sides can check: response and resolution times per priority, the hours the clock runs, what pauses it, any uptime target, and what happens when a target is missed.
An SLA is the promise to the customer. An OLA (operational level agreement) is an internal agreement between teams that makes the SLA possible, such as the network team committing to pick up an escalation within one hour so the service desk can meet a four-hour fix. Commitments from outside suppliers are covered by underpinning contracts.
Response time is how long until a person acknowledges the ticket and starts work. Resolution time is how long until the problem is fixed and the user confirms it. Response is fully in the provider's control, so it is the easier promise; resolution targets usually list exceptions such as waiting on the customer or a third party.
Whatever the SLA says. The most common remedy is a service credit, a discount on the next invoice. For example, the Amazon EC2 SLA (updated May 2022) gives a 10% credit when monthly uptime drops below 99.99%, 30% below 99.0% and 100% below 95.0%. Service desk SLAs may instead trigger a review meeting or, after repeated breaches, an exit clause.

MSP AI Agents

Yes. In production MSP shops today, 10% to 25% of tickets close before a human opens them. Thread alone has processed 173 million tickets across 750-plus MSP partners at 96% triage accuracy, handing back 490,000-plus technician hours. Agents own the low-risk, high-volume work (password resets, MFA enrollment, known installs, onboarding and offboarding) and flag anything that touches production data or needs judgment for a human to take.
On a five-person desk, reported deployments show $78,000 to $130,000 in annual direct labor savings, roughly 30% fewer escalations, and 15% to 20% better SLA compliance. Broader MSP adoption data adds ticket handling time cut by 45% and five to 12 points of margin, all from reclaimed capacity rather than headcount cuts.

About OpenFrame

OpenFrame isn't built to plug into your stack. It replaces it. Instead of duct-taping a dozen tools together (RMM, MDM, SIEM, patching, remote access, each its own login and bill), we bundle it into one unified platform: RMM, MDM, monitoring, automation, remote access, patch management, security monitoring, and ticketing, plus built-in AI copilots. So "does it integrate with X?" usually means: you won't need X anymore.
Most platforms give you one piece and expect you to bolt the rest on. OpenFrame unifies the whole stack in one place, with AI copilots built in. Fewer logins, fewer bills, less duct tape.
In the cloud, on US soil. Your data stays stateside.
Both. It's built for MSPs and MSSPs alike.