The scoreboard, honestly: the hard targets, how often each one is actually looked at,
and the quiet human signals that never make it onto a dashboard.
Incident Response SLA Adherence
How quickly you respond to and resolve incidents that come your way, especially the urgent ones.
Target · >95% of P1/P2 incidents responded to within 15 minutes and resolved within 2 hours.You pick up a P1 alert for a database outage at 10:03 AM, acknowledge it by 10:05 AM, and have the database back online by 10:45 AM, well within the 2-hour target.
System Availability for Owned Services
The actual uptime of the specific infrastructure components or services you're responsible for.
Target · Maintain 99.9% ('three nines') uptime for your assigned systems (e.g., vSphere cluster, backup solution).Your assigned virtualisation cluster had 8 hours of downtime over the last year, which translates to roughly 99.908% availability, just hitting the target.
Patching Compliance Rate
Ensuring critical security patches are applied to systems you manage within the agreed timeframe.
Target · 98% of critical patches deployed within 72 hours of release to production environments.Out of 50 critical patches released last month, you successfully deployed 49 within 72 hours, meaning a 98% compliance rate.
Documentation Quality and Completeness
How well you document new configurations, changes, and troubleshooting steps for systems you own.
Target · All new system deployments or major changes have associated runbooks and architecture diagrams completed within 5 working days.After deploying a new application server, you've updated the CMDB entry, created a basic runbook for common issues, and added it to the monitoring dashboard, all within the week.
Proactive Problem Identification
You're not just reacting to alerts; you're spotting potential issues before they become outages. This means looking at trends, not just thresholds.
- You'll bring potential issues to your manager's attention (e.g., 'I've noticed our disk I/O on the SQL server is slowly creeping up, we should look into it'). You'll propose small changes to prevent future problems. Your colleagues will say you're good at spotting things early.
Effective Troubleshooting Behaviour
When things break, you approach the problem methodically, gather facts, and communicate clearly, rather than panicking or jumping to conclusions.
- During an incident, you'll provide clear, concise updates. You'll follow established runbooks where available, and document your steps when you're in uncharted territory. Your post-incident reviews will show a logical progression of investigation and resolution.
Collaboration with Development Teams
You work well with the software development teams, helping them understand infrastructure needs and offering solutions to their deployment challenges.
- Dev teams will come to you for advice on how to best configure their applications for our infrastructure. You'll help them diagnose issues that cross the application/infrastructure boundary. You'll be seen as helpful and approachable, not a blocker.
Contribution to Team Knowledge & Best Practices
You share what you learn, improve existing processes, and help the team get better at what we do.
- You'll update existing runbooks with new information. You'll contribute to our internal knowledge base. You might even offer to do a quick 'lunch and learn' session on a new tool or technique you've picked up. You'll offer constructive feedback during code reviews for IaC.