Machine Learning Ops Engineer
We are investing in in-house LLM tooling and are hiring a dedicated MLOps Engineer to help grow it. You will build AI-powered capabilities—retrieval-augmented generation, tool integrations, and agentic workflows—and turn them into reliable services used by teams across the company. This is a builder's role focused on shipping new capability.
The role spans a broad stack. We welcome both generalists and specialists—you do not need every skill listed below. Tell us where you are strong and where you want to grow. The center of gravity is LLM application development, retrieval quality, and agent design.
Responsibilities:
LLM Applications, RAG & Agents
Design and build new LLM-powered tools and agentic workflows that automate real work and improve productivity across the company
Extend and improve our RAG systems—ingestion, chunking, embedding, retrieval, ranking, and evaluation—to raise answer quality
Structure retrieval around the organization's information hierarchy so that relevance and access boundaries improve together
Build tool integrations that connect LLMs to internal systems and data sources
Design agents that act safely against real systems, with appropriate guardrails, human-in-the-loop where warranted, and clear failure behavior
Establish evaluation and testing frameworks to measure quality, catch regressions, and guide iteration
Partner with teams across the company to identify high-value use cases and turn them into deployed tools
Service Deployment & AI Infrastructure
Deploy AI tools and services for teams across the company, taking them from prototype to reliable production
Build and operate the infrastructure that hosts models, tools, and supporting services on Kubernetes
Manage model serving, inference endpoints, and the APIs and gateways around them
Implement monitoring, logging, and usage observability so we understand how tools perform and get used
Access, Security & Data Boundaries
Ensure retrieval and agent tools respect the same access boundaries as the underlying systems—no cross-team or cross-project data leakage
Integrate with existing identity and permission systems so tools honor who is allowed to see what
Apply data-handling practices appropriate to a defense environment
Treat access control as a first-class design concern in every tool, not an afterthought
Automation & Data Operations
Build CI/CD pipelines for AI tools, services, and agents
Automate provisioning and configuration with Ansible and infrastructure-as-code practices
Build data pipelines to ingest, transform, and index content for RAG and AI applications
Manage vector databases and other stores backing retrieval and AI workloads, including versioning and quality checks
Maintain reproducible environments across development, staging, and production
Qualifications:
Bachelor's in Computer Science, Software Engineering, Data Engineering, or related field – equivalent industry experience also welcome
3-6+ years of experience in MLOps, software, platform, or backend engineering (relevant depth matters more than exact years)
Strong proficiency in Python and comfort building, shipping, and operating services
Experience building LLM-powered applications—working with LLM APIs or self-hosted models, prompts, and tool/function calling
Hands-on experience with Kubernetes and containerized deployment
Solid understanding of CI/CD, infrastructure-as-code, and production service reliability
Awareness of access control and data-boundary concerns when connecting tools to sensitive internal systems
Demonstrated ability to learn quickly and work across unfamiliar parts of the stack
Depth in at least one core area—LLM application development, RAG/retrieval, agent design, or AI infrastructure—with genuine interest in growing into the others
Preferred:
Hands-on experience with RAG systems, embeddings, and vector databases (pgvector, Qdrant, Weaviate, Milvus, or similar)
Experience designing and shipping agentic workflows, including tool use, orchestration, and guardrails
Familiarity with the Model Context Protocol (MCP) or similar tool-integration frameworks for LLMs
Experience integrating LLM tools with enterprise systems (productivity suites, business systems, or developer platforms) via their APIs
Knowledge of LLM evaluation, prompt engineering, and quality/regression measurement
Experience serving models and optimizing inference (vLLM, TGI, Triton, or similar)
Familiarity with agent/orchestration libraries (LangChain, LlamaIndex, or equivalent)
Experience with Ansible for configuration management and automation
Experience implementing identity, authentication, and fine-grained authorization (OAuth, SSO, RBAC)
Observability experience for AI/ML workloads, including usage and quality metrics
GPU infrastructure and scheduling experience for training or inference
Understanding of security and data-handling requirements in regulated or defense environments
Ability to obtain or maintain a security clearance
Pay range for this role
$140,000—$175,000 USD