Senior Data Engineer
ABOUT THE ROLE
As a Senior Data Engineer, you will play a key role in building and scaling the data platform that powers analytics across Trust Wallet. You'll be responsible for designing, developing and maintaining data pipelines and data models, ensuring a seamless, scalable and reliable flow of data from source systems through to the decisions it supports. Leveraging your experience with Databricks and Spark across both streaming and batch workloads, dbt or a similar transformation framework, cloud infrastructure, and custom solutions in Python, you'll work closely with our data analysts and engineering teams to turn data into a strategic asset for our company.
KEY RESPONSIBILITIES
- Data Platform Engineering: Architect and maintain robust, scalable and secure data infrastructure on Databricks, covering both streaming and batch workloads.
- Data Pipeline Development: Design, develop and maintain data pipelines, primarily in Python and Spark, to automate ingestion and transformation across internal systems, external providers and on-chain sources.
- Data Modelling: Design and maintain dimensional data models and transformation layers in dbt, with tests and documented contracts, so metrics are consistent and reusable across the company.
- Data Lake Management: Oversee the data lake and lakehouse layers, ensuring efficient storage, effective partitioning, high data quality, and monitoring and alerting that surfaces issues early.
- Integration and Customisation: Integrate Databricks with a wide range of data sources, including change data capture from operational databases, third-party APIs and blockchain data, and adapt data flows to specific business needs.
- Performance, Scalability and Cost: Optimise pipelines and storage for performance, reliability and cost efficiency at scale.
- Data Governance and Security: Apply best practices for governance, security and compliance in cloud and Databricks environments, including access control, encryption and monitoring.
- Collaboration and Documentation: Work closely with platform engineers, data analysts and other stakeholders to understand data requirements, and document infrastructure, models and best practices.
REQUIRED QUALIFICATIONS
- Experience in Data Engineering: 3+ years as a Data Engineer, with hands-on ownership of production pipelines and lakehouse or data warehouse architecture.
- Databricks and Spark: Strong practical experience with Databricks, Delta Lake and Spark, including both streaming and batch processing, and the judgement to choose between them.
- Data Modelling: Solid understanding of dimensional modelling, slowly changing dimensions and data warehouse design, with the ability to define and defend the grain of the models you build.
- Transformation Frameworks: Experience with dbt or an equivalent framework for managing transformations, testing and lineage.
- Cloud Proficiency: Strong cloud fundamentals, including identity and access management, object storage, networking and infrastructure as code. We run on AWS; deep experience with Azure or GCP transfers well.
- Proficiency in Python and SQL: Comfortable writing production transformations as well as services and connectors.
- Data Quality and Governance: Experience implementing testing, monitoring and governance practices in cloud data environments.
PREFERRED SKILLS
- Experience with change data capture and replication from operational databases
- Familiarity with blockchain or on-chain data
- Infrastructure as code, ideally Terraform, and CI/CD for data workloads
- Experience with BI and semantic layers such as Holistics, Looker or similar
- Containerisation and orchestration (Docker, Kubernetes)