The pathway
How you actually get there, here
How you become one varies far more by country than what one does. This is the UK route. Most people take one of these ways in; the right one depends on where you're starting from.
- 1
From Data Engineer (L2)
2-3 years as a Mid-level Data EngineerSkills to master
- Independent ownership of complex pipelines, strong data modelling skills, proactive problem-solving, and initial informal mentorship of junior colleagues.
You're ready to move on when
- Consistently delivering reliable data solutions without close supervision.
- Proactively identifying and resolving data quality issues before they escalate.
- Actively participating in design discussions and offering valuable technical insights.
- Being the 'go-to' person for specific data domains or technical challenges within the team.
- 2
From Software Engineer with Data Focus
3-5 years as a Software Engineer, with 1-2 years specifically on data-heavy projectsSkills to master
- Deepening knowledge of distributed systems (Spark, Kafka), dimensional modelling, and specific cloud data services (AWS Glue, Redshift, Snowflake). Understanding the nuances of data quality and governance.
You're ready to move on when
- Strong software engineering fundamentals (clean code, testing, CI/CD).
- Experience building robust, scalable backend systems that handle large volumes of data.
- A clear interest and some experience in data transformation, warehousing, or analytics.
- A willingness to learn the specific data modelling and platform tools we use.
- 3
From Data Analyst/Scientist with Strong Engineering Skills
4-6 years as a Data Analyst/Scientist, with significant time spent on data preparation and pipeline buildingSkills to master
- Transitioning from data consumption/analysis to data *production* and platform building. This means focusing on infrastructure as code, distributed computing, and robust pipeline orchestration.
You're ready to move on when
- Consistently building and maintaining complex SQL queries and Python scripts for data extraction and transformation.
- Frustration with existing data quality or pipeline reliability, and a desire to fix it at the source.
- A solid understanding of data needs from a consumer perspective, coupled with a growing interest in the underlying engineering.
- Some experience with cloud platforms (AWS, GCP, Azure) and version control (Git).