The pathway
How you actually get there, here
How you become one varies far more by country than what one does. This is the UK route. Most people take one of these ways in; the right one depends on where you're starting from.
- 1
Model Reliability Engineer (L2)
2-3 years as an MRESkills to master
- Independent incident resolution, building comprehensive monitoring, automating routine deployments, understanding model lifecycle.
You're ready to move on when
- Consistently resolving P2 incidents without escalation.
- Successfully implementing new monitoring for several models.
- Proactively identifying and fixing potential reliability issues.
- Mentoring new joiners informally and providing solid code reviews.
- 2
Senior Site Reliability Engineer (SRE) or DevOps Engineer
5+ years in SRE/DevOpsSkills to master
- Deep understanding of ML-specific challenges (drift, skew, model serving), Python programming for ML, MLOps tooling.
You're ready to move on when
- Strong background in distributed systems, cloud infrastructure, and CI/CD.
- Demonstrable interest and some experience with machine learning concepts or platforms.
- Ability to quickly pick up new domains and apply SRE principles to ML.
- Proven track record of building and operating highly available systems.
- 3
Senior Data Engineer or ML Engineer
5+ years in Data/ML EngineeringSkills to master
- Deep dive into production operations, incident management, infrastructure as code, advanced monitoring, and system resilience.
You're ready to move on when
- Strong programming skills and understanding of data pipelines.
- Experience deploying and managing ML models in production (even if not full MLOps).
- Desire to specialise in operational excellence and system reliability.
- Comfortable with on-call rotations and incident response.