The scoreboard, honestly: the hard targets, how often each one is actually looked at,
and the quiet human signals that never make it onto a dashboard.
Mean Time To Resolution (MTTR) for P1/P2 Incidents
The average time it takes to restore service for our most critical incidents.
Target · < 30 minutes for P1, < 2 hours for P2If a critical payment gateway goes down (P1), we'd expect it to be back up within 30 minutes, or you'll be explaining why. For a P2, like a non-critical reporting tool, we're aiming for under two hours.
Incident Recurrence Rate
The percentage of incidents that reoccur within a 30-day window after being 'resolved'.
Target · < 5% for P1/P2 incidentsIf we fix a bug that causes a P2 outage, but it pops up again two weeks later because the root cause wasn't fully addressed, that counts against you. We're aiming for lasting fixes, not just quick workarounds.
Root Cause Analysis (RCA) Quality Score
Evaluates the thoroughness, accuracy, and actionability of post-incident RCAs, including identified preventative measures.
Target · > 90% score on internal auditYour RCA for the Q2 database outage should clearly identify the exact misconfiguration, detail the impact, and propose concrete, measurable actions to prevent it ever happening again. If it's vague or misses key details, it won't pass muster.
Team SLA Adherence
The percentage of service requests and incidents handled by your direct reports that meet their defined Service Level Agreements.
Target · > 95% across all prioritiesIf your team has a target to resolve P3 tickets within 8 hours, and they consistently hit 90%, you'll need to figure out why and get them back on track. This reflects on your leadership and process design.
Process Improvement Savings/Efficiency Gains
Quantifiable benefits (time saved, cost reduced, error rate decreased) from process improvements you've designed and implemented.
Target · Identify and deliver £50K in annualised savings or 10% efficiency gainYou might redesign the change management process, reducing manual approval steps and saving 50 hours of staff time per month across the department—that's a clear win. Or maybe you automate a reporting task that used to take three days, freeing up an analyst for more complex work.
Incident Communication Clarity & Timeliness
How effectively and promptly you communicate during major incidents to all relevant internal and external parties.
- Feedback from executive leadership and affected business units
- absence of follow-up questions due to unclear updates
- consistent use of agreed communication channels and templates
- proactive updates during extended outages.
Team Development & Morale
The growth of your direct reports and the overall health of your team.
- Retention rates within your team
- feedback from 1:1s and skip-level meetings
- successful completion of development plans by team members
- positive peer feedback on your mentoring
- your team members taking on more complex tasks confidently.
Stakeholder Confidence & Trust
The level of trust and reliance key internal and external partners place in your ability to manage service delivery.
- Being proactively consulted on new product launches or system changes
- positive feedback from Product, Engineering, and Customer Success teams
- stakeholders coming to you directly for advice on operational challenges
- minimal escalations above your head during incidents.
Proactive Problem Identification & Resolution
Your ability to spot recurring issues, underlying weaknesses, or potential future problems before they become major incidents.
- Number of 'problem tickets' you initiate that lead to significant service improvements
- early detection of trends from monitoring data
- successful implementation of preventative measures that avoid predicted outages
- a reduction in 'surprise' incidents.