The scoreboard, honestly: the hard targets, how often each one is actually looked at,
and the quiet human signals that never make it onto a dashboard.
Mean Time To Recovery (MTTR) for Owned Services
How quickly we get a critical service back online after an incident, specifically for the services you and your team are responsible for.
Target · Drive a 25% reduction in MTTR for your key services over 6 months.If a service you own typically takes 60 minutes to recover, you'd aim to get that down to 45 minutes by implementing better runbooks, automation, or monitoring.
Toil Reduction for Your Team
The amount of repetitive, manual operational work that your team eliminates through automation or process improvements.
Target · Automate away >5 hours per week of manual tasks for each engineer in your team.If your team spends 10 hours a week on manual certificate renewals, you'd aim to automate 5 hours of that, freeing up time for more impactful work.
Alert Signal-to-Noise Ratio
The percentage of alerts that are genuinely actionable and require human intervention, versus those that are just noise.
Target · Reduce non-actionable alerts by 50% across your owned services.If your service generates 100 alerts a week, and 70 of them are false positives or non-critical, you'd aim to get that down to 35 non-actionable alerts.
Cloud Cost Optimisation for Owned Services
Identifying and implementing ways to reduce our cloud spend without compromising performance or reliability.
Target · Achieve a 10-15% cost reduction for your primary services annually.Moving a specific workload from on-demand EC2 instances to reserved instances, or identifying and shutting down idle development environments, saving £2,000 per month.
Architectural Robustness & Resilience
How well your designs and implementations improve the overall stability and fault tolerance of our systems.
- Reduced incidence of specific failure modes
- successful disaster recovery drills
- positive feedback from development teams on system stability
- clear, well-documented architectural decisions.
Mentorship & Team Development
Your ability to guide, teach, and uplift the engineers in your team, helping them grow their technical skills and operational maturity.
- Junior engineers taking on more complex tasks
- positive feedback in 1:1s and performance reviews
- successful delegation of significant work
- team members actively seeking your advice and guidance.
Incident Leadership & Postmortem Quality
Your effectiveness in leading critical incidents, making sound decisions under pressure, and driving thorough, blameless postmortems.
- Swift resolution of P1/P2 incidents you lead
- comprehensive, actionable postmortem reports
- identified root causes leading to preventative measures
- calm and clear communication during incidents.
Documentation & Knowledge Sharing
The clarity, completeness, and maintainability of the documentation and runbooks you and your team produce.
- New team members can effectively use your documentation
- reduced questions about common procedures
- up-to-date Confluence pages
- successful handovers of new systems.