Sr. Machine Learning Engineer
About the Role:
We’re hiring a Sr. Machine Learning Engineer to join the ML Engineering team. Our team builds the production systems that make 6sense’s machine learning dependable at massive scale: model-training and refresh pipelines, evaluation and release controls, high-throughput batch and online inference, and the shared platforms that let Data Science and product teams ship safely.
This is a hands-on engineering role for someone who enjoys owning the hard middle between a strong model and a reliable customer capability. You will work closely with Data Scientists, platform engineers, and product teams to turn experimentation into governed, observable, scalable production systems. You will help shape how models and AI agents are evaluated, promoted, monitored, and improved—not simply deploy them once.
What You’ll Do :
Own production ML capabilities end to end: turn a business or modeling need into a well-designed pipeline or service, then operate and improve it in production.
Build and evolve scalable training, model-refresh, feature/data-validation, and inference workflows across batch and real-time use cases.
Design reliable model lifecycle controls: experiment tracking, evaluation gates, model/version promotion, rollback, lineage, reproducibility, and monitoring.
Build platform primitives that enable Data Science and AI product teams to ship faster without compromising reliability, security, or cost.
Improve the performance, resilience, and observability of distributed ML workloads and model-serving systems.
Partner with Data Science on evaluation design, data quality, model-health signals, and production debugging.
Contribute to LLM/agent evaluation and serving infrastructure where appropriate, including offline and online evaluation, tracing, quality gates, and regression detection.
Lead technical design for ambiguous projects, influence architecture across teams, and mentor engineers through code reviews and hands-on guidance.
Communicate decisions, risks, and operational status clearly to engineering, product, and leadership stakeholders.
What We’re Looking For :
Required :
6+ years of industry experience building and operating production machine-learning or data-intensive distributed systems, including substantial end-to-end ownership.
Strong Python engineering skills and practical experience designing maintainable, testable services and pipelines.
Demonstrated MLOps depth: experiment tracking, model registry/versioning, CI/CD, reproducible training, data/model validation, deployment strategies, rollback, and production monitoring.
Experience with distributed data and ML infrastructure such as Spark, Ray, Databricks, Kubernetes, AWS, or equivalent platforms.
Strong understanding of model-training and inference trade-offs: data quality, feature engineering, evaluation, latency/throughput, cost, reliability, and model drift.
Experience productionizing at least one of: classical ML models, deep-learning/NLP models, embedding/retrieval systems, or LLM/agent workflows.
Solid judgment in incident response and operational ownership; able to diagnose failures across data, model, infrastructure, and serving layers.
Ability to translate ambiguous product and Data Science requirements into a pragmatic technical plan and drive it to completion.
Clear written and verbal communication with both technical and non-technical partners.
Nice to Have :
Hands-on experience with MLflow, Databricks, Ray, Kubernetes, Triton/managed model serving, or similar ML platform tooling.
Experience operating high-volume batch scoring or low-latency online inference systems.
Experience with LLM/agent evaluation frameworks, tracing/observability, RAG, vector search, LangGraph/LangSmith, or Amazon Bedrock.
Experience with feature stores, data contracts, schema validation, and data-quality systems.
Experience in B2B SaaS or a high-scale data platform where reliability and customer impact matter.