Middle Data Engineer
Responsibilities:
Build, maintain, and optimise scalable ETL/ELT pipelines (batch and near-real-time) on Azure Data Cloud Platform (e.g., Data Lake, Microsoft Fabric, Azure Data Factory).
Develop and refine data models to support BI reporting, analytics, and ML/AI use cases.
Write efficient, well-documented T-SQL and PySpark code following team coding standards.
Implement automated testing, data validation, and monitoring (SLAs, alerts) to ensure pipeline reliability.
Contribute to data governance practices, including lineage tracking, metadata management, and quality controls.
Support CI/CD pipelines for data assets, ensuring version control and reproducibility.
Partner with analytics engineers to scope, refine, and prioritise data requirements from business stakeholders.
Work with Analysts, BI Developers, Data Scientists, and business teams to translate requirements into production-ready data solutions.
Provide input on data readiness for machine learning and analytics projects.
Contribute to the evolution of the ED&I data platform, including tooling, standards, and documentation.
Stay current with emerging data engineering patterns and technologies; propose improvements to team processes.
Leverage AI-driven development tools (e.g., generative-AI code assistants, automated data profiling) to accelerate delivery.
Support performance tuning and cost optimisation across the data platform.
Requirements:
3+ years in data engineering or a closely related role.
Bachelor’s degree in Computer Science, Data Engineering, or a related field.
Strong T-SQL skills and working proficiency in PySpark or Python for data processing.
Hands-on experience with MS Azure Storage Explorer and SSMS.
Hands-on experience with cloud-based data engineering services and orchestration tools (e.g., Azure Data Factory, Microsoft Fabric).
Practical experience building ETL/ELT pipelines and dimensional or analytical data models.
Familiarity with CI/CD practices in data engineering, including version control (Git) and automated testing.
Nice to have:
Experience with real-time or streaming data architectures.
Experience with PowerShell, Apache Kafka, and/or KQL.
Exposure to AI/ML workflows (feature engineering, data preparation for model training).
Familiarity with Power BI or other BI/visualisation tools.
Experience using AI productivity tools (e.g., ChatGPT, Claude, Copilot, Cursor) in day-to-day and data engineering tasks.
Understanding of data security, privacy, and compliance considerations.