Senior AI Engineer
We're looking for a Senior AI Engineer to help build and scale dunnhumby's Enterprise AI Platform- designing, deploying, and operating production grade AI systems used across engineering teams. You'll work across the full AI lifecycle: model training and fine-tuning, agentic workflows, RAG, AI observability, and AI-powered user experiences, using the latest advancements in Generative AI.
Key Responsibilities
Build reusable, scalable AI services for prompt orchestration, model routing, embeddings, structured generation, and tool calling; develop configurable multi-provider AI runtimes and secure cloud-native microservices.
Design multi-agent and autonomous systems with reasoning, planning, memory, and tool execution; build graph-based, long-running workflows with human-in-the-loop checkpoints using MCP and A2A.
Build enterprise-grade RAG pipelines— ingestion, chunking, embeddings, hybrid search, reranking, citations — and continuously evaluate retrieval quality.
Train, fine tune (LoRA/QLoRA/PEFT), and evaluate ML/DL models; build training pipelines, run experimentation and hyperparameter optimization, and productionize models with data science partners.
Deploy, monitor, and continuously improve agents and models in production — experiment tracking, model registry, versioning/rollback, drift and cost monitoring, CI/CD, and canary/blue-green deployments.
Implement guardrails for hallucination, prompt injection, and PII; establish evaluation, monitoring, and responsible-AI compliance practices.
Build and deploy cloud-native AI services (Docker, Kubernetes, Terraform) on GCP and Azure with autoscaling, observability, and distributed tracing; own services from build through production support.
Build responsive React-based interfaces for chat, copilots, prompt playgrounds, and agent/evaluation dashboards, integrated via REST, SSE, and WebSockets.
Write clean, tested code; drive architecture reviews, code reviews, and mentor engineers.
Required Qualifications
Bachelor's/Master's in Computer Science, AI, Engineering, or related field.
8+ years of software engineering experience, including 3+ years building production AI/ML applications.
Strong grounding in distributed systems, cloud-native architecture, and microservices.
Track record delivering enterprise-grade AI solutions from POC to production.
Technical Skills
Programming: Python (expert), TypeScript/JavaScript, SQL, async programming, REST & gRPC, design patterns
AI/ML Concepts: LLMs, prompt engineering, embeddings, RAG, hybrid search, agentic AI (tool/function calling, MCP, A2A), model training & fine-tuning (LoRA/QLoRA/PEFT), hyperparameter optimization, GPU optimization, quantization
AI Frameworks: LangChain, LangGraph, Google ADK, OpenAI Agents SDK, CrewAI, AutoGen, LlamaIndex, Semantic Kernel, DSPy, Pydantic AI
Cloud AI Platforms: Google Vertex AI, Azure AI Foundry, Azure OpenAI, Model Garden
ML & MLOps: PyTorch, TensorFlow, Hugging Face, MLflow, Kubeflow, model registry, experiment tracking, continuous training
AgentOps & Observability: LangSmith/Langfuse/Arize Phoenix, OpenTelemetry, Grafana, New Relic, agent evaluation, cost monitoring
Vector Databases: Pinecone/Weaviate/Milvus/pgvector/Vertex AI Vector Search/Azure AI Search
Cloud & DevOps: Docker, Kubernetes, Terraform, ArgoCD, GitHub Actions/Azure DevOps, CI/CD, IaC
Workflow Orchestration: Temporal/Argo Workflows/event-driven architecture/message queues