The pathway
How you actually get there, here
How you become one varies far more by country than what one does. This is the UK route. Most people take one of these ways in; the right one depends on where you're starting from.
- 1
Mid-level Chaos Engineer
2-3 yearsSkills to master
- Independent execution of GameDays, strong blast radius analysis, basic fault injection techniques, effective use of observability tools.
You're ready to move on when
- Consistently delivers well-documented experiment reports.
- Can debug minor issues during experiments without supervision.
- Proactively identifies opportunities for simple chaos experiments.
- Has a solid understanding of our core services and their dependencies.
- 2
Senior Site Reliability Engineer (SRE)
3-5 yearsSkills to master
- Deep understanding of system internals, incident response leadership, automation of operational tasks, strong troubleshooting skills in production environments.
You're ready to move on when
- Has led multiple high-severity incident responses.
- Demonstrates a proactive approach to improving system reliability.
- Strong coding skills for automation and tooling.
- Excellent understanding of our production environment and infrastructure.
- 3
Senior Software Engineer (Backend/Distributed Systems)
4-6 yearsSkills to master
- Deep expertise in building and operating distributed systems, strong coding and architectural design skills, understanding of performance and scalability challenges.
You're ready to move on when
- Has designed and implemented complex backend services.
- Understands common failure modes in distributed systems from a development perspective.
- Can write robust, testable, and resilient code.
- Shows a keen interest in system reliability and proactive failure testing.