SLAs are not the enemy

Miami University

No reason to arm yourself against the zombie apocalypse – the shared responsibility for logging incidents should not be a scary event. However, there is some information you will need to be armed with, to be successful!

One major deliberate change in our new Incident Management process is that all incidents will no longer have the same priority.

In the old process, solvers are left to decide which incidents should be addressed in which order: first-in, first-out; cherry-picking; favorite-client first; etc. The new process puts the incident priority into the drivers seat. The priority of the ticket, calculated by impact and urgency, determines the resolution timeline.

At ticket logging time, the logger selects impact and urgency based on the following rubric. The project team is still refining this as we test it in practical settings.

The combination of impact and urgency result in a priority designation (1 – 5) with two time-frames attached:

  1. time-to-act: a commitment to have human eyes on the ticket by the Support Desk
  2. time-to-complete:  the time the client is back in business; this is not necessarily the same as repairing the service!

P1

P2

P3

P4

P5

“SLA”

24×7

24×7

24×7

business hours

business hours

time-to-act (hours)

.5

1

1

1

1

time-to-complete (hours)

2

4

24

16 (2 bus. days)

56 (7 bus. days)

In the new process, all incidents will be reported against these targets, which has several implications for you!

  • Solvers (those to whom unsolved incidents are escalated) will use the priority and SLA deadline to drive the order in which you address incidents. We anticipate that checking your incident queue twice daily should be sufficient to stay on top of them

  • Priority 1 incidents trigger a special major incident process to bring immediate attention to the incident

  • Priority 2 incidents trigger the major incident process once the 4 hour window has been breached — again to bring immediate attention to the situation.

  • There will be times when we fall short and fail to complete an incident before the target is breached. These instances will provide an opportunity to investigate the cause and make incremental improvements to the process, training, tools, or documentation at the heart of the breach.

6 responses to “SLAs are not the enemy”

  1. wardtd Avatar
    wardtd

    Question about time-to-act. As there’s not a good way to prioritize incoming incidents/requests before seeing them (we can’t know whether it’s a priority 1 or a priority 5 before we’ve read their email), how do we handle the fact that there’s a different target for different priority levels? It seems like we’d need to get eyes on all tickets within .5 hours to ensure we were hitting that target. Is time-to-act referring to that initial classification by level 1 staff, or does it refer to a guaranteed time at which work will begin after it has been escalated beyond level 1?

    1. blackrw Avatar
      blackrw

      Sorry I missed your comment earlier Tim. Time to act is the expectation is that a human being should have eyes on it within an hour (business hours for Priority 4&5, 24×7 for P1-3).

      The team felt that we need to be clear with the university community that phone is the best way to report an urgent situation. The 30 minute timer on P1 is simply an allowance for Wright State to get hold of someone during overnight hours.

      It’s important to remember that these are targets — the team wishes they weren’t called SLAs, but that’s what the tool calls them. If we find a particular gap, then this data helps make the case to improve the process, tool, or resources to achieve them (or lower the targets).

  2. blackrw Avatar
    blackrw

    To your actual question: my first reaction is that so long as that we are meeting those targets (no breaches) universally and consistently, then I’m not sure the order matters. I suppose if that’s the case — we should consider looking at increasing the target? Thoughts?

    To your non-question: The team is somewhat conflicted about ever calling these targets SLAs in the first place. The SLA should be at a much higher level and your point is well taken: these should be formally negotiated with the decision-makers. As we look to deploy a service-level-management process down the road, we hope to use that opportunity to more formally vet these expectations.

    1. blackrw Avatar
      blackrw

      This comment was supposed to be a reply to Mike. I fail at blog commenting.

  3. beckmd Avatar
    beckmd

    Aside: My favorite part about working on SLAs was when I was asked to remove the client signature line from a draft agreement, as it was not something that was actually going to be agreed upon.

    Actual question: It seems like many typical incidents will have the same priority. What can be done to prevent FIFO, cherry-picking, etc.?

    1. blackrw Avatar
      blackrw

      To your actual question: my first reaction is that so long as that we are meeting those targets (no breaches) universally and consistently, then I’m not sure the order matters. I suppose if that’s the case — we should consider looking at increasing the target? Thoughts?

      To your non-question: The team is somewhat conflicted about ever calling these targets SLAs in the first place. The SLA should be at a much higher level and your point is well taken: these should be formally negotiated with the decision-makers. As we look to deploy a service-level-management process down the road, we hope to use that opportunity to more formally vet these expectations.

Leave a Reply

Your email address will not be published. Required fields are marked *