Senior Data Engineer - Internal Data Platform & Analytics
About the Role:
We are looking for an individual-contributor Data Engineer to help build and operate Glean’s internal data platform. This is an internal-facing data platform role focused on analytics engineering for Glean’s internal teams and not customer-facing implementation. The role will focus on reliable, governed, cost-efficient data infrastructure and analytics foundations used by internal teams across the company. You will work across data ingestion, transformation, modeling, quality, access governance, and platform reliability. This role is hands-on and well suited to an engineer who enjoys turning ambiguous data problems into durable systems and clear operating standards. This is an in-person role based in Bangalore, India, with regular office presence expected.
You will:
Build and maintain reliable batch and API-based data ingestion pipelines.
Improve data quality through testing, continuous integration (CI) checks, ownership metadata, and clear layer boundaries.
Operate and improve BigQuery data infrastructure with an emphasis on performance and cost efficiency.
Implement data access controls, governance workflows, and safe self-serve access patterns.
Improve pipeline observability, failure classification, incident triage, and recovery processes.
Partner with Data Science, Business Intelligence, Finance, Sales Operations, Marketing, Security, Reliability Engineering, and other internal teams to understand data needs and deliver reusable platform capabilities.
Participate in design reviews, code reviews, documentation, and operational support for the data platform.
About you:
Minimum experience: 7–10 years overall, including at least 7 years of data engineering experience.
Strong Data engineering fundamentals and experience building production data systems.
An exceptionally high AI proficiency through habitual, high-value use of LLMs; sound judgment about when and how to apply them; rigorous validation and workflow improvement
Experience with SQL and Python, or comparable programming languages.
Experience with a cloud data warehouse, preferably BigQuery or a similar platform.
Experience with data transformation frameworks such as DBT, including testing and deployment workflows.
In Depth Understanding of Columnar File systems like parquet, Hudi Or Iceberg.
Understanding of dimensional modeling, data contracts, lineage, and data quality practices.
Experience designing or operating APIs, batch pipelines, or event-driven ingestion systems.
Ability to communicate technical trade-offs clearly and work effectively with internal stakeholders.
Ownership mindset: you can take a problem from discovery through implementation, rollout, and operational follow-through.
Ability to maintain a productive collaboration between IST and US PST time zones
Experience with BigQuery governance, IAM/RBAC, policy tags, masking, streaming systems or cost controls.
Experience building reusable data platform frameworks rather than one-off pipelines.
Familiarity with semantic layers, metric stores, or systems that make trusted data consumable by AI and analytics tools.
Experience with data observability, orchestration, CI/CD, or infrastructure-as-code.
Experience working in a fast-growing company where requirements and priorities evolve quickly.
Location:
This role is hybrid (4 days a week in our Bangalore office)
Compensation & Benefits:
Compensation offered will be determined by factors such as location, level, job-related knowledge, skills, and experience. Certain roles may be eligible for variable compensation, equity, and benefits.
We’re committed to building and sustaining a diverse, inclusive workplace. We strive to attract and retain people with a wide range of backgrounds, experiences, and perspectives, and we do not discriminate on the basis of gender, ethnicity, sexual orientation, religion, civil or family status, age, disability, or race.
#LI-HYBRID