Senior AI/ML engineer
We are looking for a Senior AI Engineer with strong LLM and product engineering experience to join our AI Platform team and help ship the AI-driven features across our app — including personalized in-app content recommendations, conversational AI-powered insights, and AI-generated mobile push notifications. You will focus on the product-facing behavior of our GenAI features (relevance, quality, latency, and safety), using the platform’s judges, fine-tuning pipelines, and evaluation tools, and contributing to infrastructure work (Databricks, AWS, Terraform) as initiatives require it.
What you’ll do
Feature development: build and ship LLM-powered product features, including AI-generated mobile push notifications, personalized in-app content recommendations and custom user journeys, and conversational AI-powered insights — including structured output such as charts and widgets
Latency and UX: reduce perceived and actual response latency through streaming, pregeneration, and inference-time tradeoffs, working closely with product and design on how AI output is experienced by users
Prompt and model behavior: iterate on prompt design and calibration, use fine-tuned health-domain models (Llama, Gemma, MedGemma) where relevant, and work with the Judge-as-a-Service ecosystem to keep outputs safe and high quality
Experimentation: help enable A/B testing for model and prompt changes, using evaluation gates to ship safely and measure impact
Data and evaluation: contribute to synthetic Q&A generation, golden test sets, and evaluation dataset design that supports the features you’re building
Infrastructure: contribute to Databricks, AWS (EKS), Terraform, or CI/CD work as initiatives require it
Cross-functional impact: collaborate with Product, Security, Analytics, and Medical teams, and engage with technology partners (Databricks, Google, OpenAI, Anthropic, AWS) on pre-release capabilities relevant to your features
Experience and skills
Must have:
7+ years of software engineering, with recent hands-on experience building product features backed by LLMs
experience with at least one of: prompt engineering, LLM evaluation, fine-tuning, or model serving
strong Python, and comfort working with APIs/SDKs that wrap LLM calls, prompt templates, and structured outputs
product sensibility: a track record of shipping user-facing features, with attention to latency, tone, and quality of AI output as experienced by end users
comfort picking up new infrastructure tools (Databricks, AWS, Terraform, GitHub Actions) as needed — deep prior expertise not required
Nice to have:
LLM evaluation frameworks (judges, graders, calibration methodology) or fine-tuning techniques (LoRA, RLHF/DPO, model distillation)
prompt optimisation frameworks (DSPy or similar)
feature stores or personalization/ranking experience (relevant to in-app content recommendation work)
healthcare, regulated-industry, or safety-critical AI systems experience
data engineering fundamentals (Spark, Delta tables, Parquet)
#LI-KP1 #LI-Hybrid
Annual Salary Range (ranges may vary based on skills and experience)
£120,000—£180,000 GBP