The scoreboard, honestly: the hard targets, how often each one is actually looked at,
and the quiet human signals that never make it onto a dashboard.
Mean Time to Resolution (MTTR) for P1/P2 Incidents
This is about how quickly we get critical services back online after a major incident kicks off. We're talking about the big, business-stopping issues.
Target · Reduce average MTTR by 15% quarter-on-quarter for P1/P2 incidents you lead.If the average P1 incident you led took 90 minutes last quarter, we'd want to see that closer to 76 minutes this quarter. It's tough, but it's where you earn your stripes.
Reduction in Recurring Incidents
This metric tells us how good you are at identifying the root cause of problems and making sure they don't keep popping up. It's about fixing the underlying issue, not just patching symptoms.
Target · Drive a 20% reduction in recurring incidents linked to your problem management efforts annually.You identify a recurring database connection issue, lead the problem investigation, and implement a fix. If that issue caused 5 incidents last year, we'd expect it to cause 4 or fewer this year, thanks to your work.
Knowledge Base Article (KBA) Contribution & Quality
You'll be expected to share your knowledge. This means writing clear, practical guides for our junior team members and even for other technical teams, helping them resolve issues faster themselves. Good knowledge articles mean less reliance on you for every little thing.
Target · Author or significantly update at least 5 impactful KBAs per quarter, with a peer review score of 4+/5.You write a step-by-step guide for troubleshooting a common application error, including screenshots. It gets used by the L1 team, reducing escalations by 10% for that specific issue. That's impactful.
SLA Adherence for Owned Workstreams
For the specific services or incident queues you're responsible for, we expect you to keep our promises to the business. This isn't just about your tickets, but the overall performance of the services under your informal guidance.
Target · Maintain >95% SLA adherence for resolution times on P3/P4 incidents within your assigned service areas.If you're looking after the 'Customer Portal' service, and it has an SLA of 4 hours for P3 tickets, you'll ensure that 95% of those tickets are resolved within that timeframe, coordinating with the relevant teams.
Incident Leadership Effectiveness
How well do you actually lead a 'war room' call? Are you clear, calm, and decisive? Do you keep everyone focused and moving towards resolution, even when things are tense?
- Positive feedback from technical leads and business stakeholders during post-incident reviews. Observed ability to de-escalate calls and maintain control. Clear, concise incident updates that keep everyone informed without overwhelming them.
Quality of Root Cause Analysis (RCA) and Post-Mortems
It's not enough to fix the immediate problem. We need you to dig deep, find the real reason something broke, and propose solid actions to prevent it from happening again. This means thorough, blameless reviews.
- Post-mortem reports that clearly identify root causes (not just symptoms), propose actionable preventative measures, and receive approval from relevant technical and business teams. Evidence of follow-through on action items.
Mentorship and Knowledge Sharing
Are you actively helping our junior team members grow? Do you make time to explain complex issues, review their work, and offer constructive feedback? This is about building capability across the team.
- Junior team members seeking your advice. Positive feedback in 1-on-1s and peer reviews about your guidance. Observable improvement in the skills and autonomy of those you mentor. Your contributions to our Confluence knowledge base.
Process Improvement Impact
You're not just following processes; you're improving them. Can you spot a bottleneck, map out a better way, and actually get it implemented? This is about making our operations smoother and more efficient.
- New or updated process documentation (e.g., in Confluence) that demonstrably reduces manual steps or error rates. Measurable improvements in workflow efficiency or service quality resulting from your proposed changes. Recognition from peers for streamlining a particular operation.