Machine Learning Ops Engineer

Zone 5 Technologies · United States · Engineering

Posted 2026-08-06

Apply for this role →

We are investing in in-house LLM tooling and are hiring a dedicated MLOps Engineer to help grow it. You will build AI-powered capabilities—retrieval-augmented generation, tool integrations, and agentic workflows—and turn them into reliable services used by teams across the company. This is a builder's role focused on shipping new capability.

The role spans a broad stack. We welcome both generalists and specialists—you do not need every skill listed below. Tell us where you are strong and where you want to grow. The center of gravity is LLM application development, retrieval quality, and agent design.

Responsibilities:

LLM Applications, RAG & Agents

Design and build new LLM-powered tools and agentic workflows that automate real work and improve productivity across the company

Extend and improve our RAG systems—ingestion, chunking, embedding, retrieval, ranking, and evaluation—to raise answer quality

Structure retrieval around the organization's information hierarchy so that relevance and access boundaries improve together

Build tool integrations that connect LLMs to internal systems and data sources

Design agents that act safely against real systems, with appropriate guardrails, human-in-the-loop where warranted, and clear failure behavior

Establish evaluation and testing frameworks to measure quality, catch regressions, and guide iteration

Partner with teams across the company to identify high-value use cases and turn them into deployed tools

Service Deployment & AI Infrastructure

Deploy AI tools and services for teams across the company, taking them from prototype to reliable production

Build and operate the infrastructure that hosts models, tools, and supporting services on Kubernetes

Manage model serving, inference endpoints, and the APIs and gateways around them

Implement monitoring, logging, and usage observability so we understand how tools perform and get used

Access, Security & Data Boundaries

Ensure retrieval and agent tools respect the same access boundaries as the underlying systems—no cross-team or cross-project data leakage

Integrate with existing identity and permission systems so tools honor who is allowed to see what

Apply data-handling practices appropriate to a defense environment

Treat access control as a first-class design concern in every tool, not an afterthought

Automation & Data Operations

Build CI/CD pipelines for AI tools, services, and agents

Automate provisioning and configuration with Ansible and infrastructure-as-code practices

Build data pipelines to ingest, transform, and index content for RAG and AI applications

Manage vector databases and other stores backing retrieval and AI workloads, including versioning and quality checks

Maintain reproducible environments across development, staging, and production

Qualifications:

Bachelor's in Computer Science, Software Engineering, Data Engineering, or related field – equivalent industry experience also welcome

3-6+ years of experience in MLOps, software, platform, or backend engineering (relevant depth matters more than exact years)

Strong proficiency in Python and comfort building, shipping, and operating services

Experience building LLM-powered applications—working with LLM APIs or self-hosted models, prompts, and tool/function calling

Hands-on experience with Kubernetes and containerized deployment

Solid understanding of CI/CD, infrastructure-as-code, and production service reliability

Awareness of access control and data-boundary concerns when connecting tools to sensitive internal systems

Demonstrated ability to learn quickly and work across unfamiliar parts of the stack

Depth in at least one core area—LLM application development, RAG/retrieval, agent design, or AI infrastructure—with genuine interest in growing into the others

Preferred:

Hands-on experience with RAG systems, embeddings, and vector databases (pgvector, Qdrant, Weaviate, Milvus, or similar)

Experience designing and shipping agentic workflows, including tool use, orchestration, and guardrails

Familiarity with the Model Context Protocol (MCP) or similar tool-integration frameworks for LLMs

Experience integrating LLM tools with enterprise systems (productivity suites, business systems, or developer platforms) via their APIs

Knowledge of LLM evaluation, prompt engineering, and quality/regression measurement

Experience serving models and optimizing inference (vLLM, TGI, Triton, or similar)

Familiarity with agent/orchestration libraries (LangChain, LlamaIndex, or equivalent)

Experience with Ansible for configuration management and automation

Experience implementing identity, authentication, and fine-grained authorization (OAuth, SSO, RBAC)

Observability experience for AI/ML workloads, including usage and quality metrics

GPU infrastructure and scheduling experience for training or inference

Understanding of security and data-handling requirements in regulated or defense environments

Ability to obtain or maintain a security clearance

Pay range for this role

$140,000—$175,000 USD

Apply for this role →

← Back to all jobs