Infrastructure Engineer, Database

LangChain · San Francisco, CA · Engineering

Posted 2026-08-20

Apply for this role →

About the team

SmithDB is LangChain's internal database team. We're building a storage and query layer purpose-built for AI observability and evaluation. Within six months we went from idea to a production system that offers industry leading performance and scalability for agent observability data. We're a small, fast team of systems engineers tackling genuinely hard problems: storage layout, query execution, compaction, and scaling toward trillions of agent traces. We develop in Rust, run on Kubernetes, and integrate tightly with S3/GCS/Azure Blob. There are no legacy constraints; this is a greenfield system with real production load and ambitious engineering goals.

About the role

We're building a database specifically designed for AI observability and evaluation, and we need someone to own the infrastructure layer that keeps it running reliably at scale. As a Database Infra Engineer on the SmithDB team, you won't be designing the storage engine — you'll be making sure the engine never goes down, scales seamlessly as our customer base grows, and is operationally excellent across cloud environments.

What you'll do

- Own the deployment and operations of SmithDB across cloud environments — including cluster lifecycle management, blue/green and rolling upgrades, and automated failover

- Build and maintain the infrastructure tooling (Terraform, Kubernetes, Helm, or equivalent) that provisions, configures, and scales SmithDB nodes

- Own the Kubernetes infrastructure that runs our distributed database services (multi-tenant, high throughput, low latency)

- Build and improve deployment pipelines, rollout strategies, and infrastructure-as-code for the storage layer

- Drive reliability engineering efforts: incident response, postmortems, SLOs, and disaster recovery for a system operating at massive scale

- Manage capacity planning and cost efficiency — model growth, rightsize resources, and ensure SmithDB can absorb traffic spikes from our largest customers without manual intervention

- Build the CI/CD pipeline for database infrastructure changes — safe, tested, and fast promotion from dev through staging to production

- Collaborate closely with SmithDB internals engineers to translate new engine features into production-ready infrastructure changes and ensure safe, low-risk rollouts

What you'll bring

- 5+ years of experience in infrastructure, platform engineering, or SRE with hands-on

- Strong hands-on experience with Kubernetes and cloud infrastructure (AWS/GCP/Azure)

- Solid scripting/systems programming ability (Go, Python, or similar);

- Experience with infrastructure-as-code and CI/CD tooling (Terraform, Helm, ArgoCD, or similar)

- Deep familiarity with at least one major cloud provider (AWS, GCP, or Azure) and the primitives used to run stateful workloads reliably — persistent volumes, managed node groups, cloud storage, etc.

- Infrastructure-as-code fluency — you write Terraform (or Pulumi/CDK) as your primary language, not an afterthought

- Strong operational instincts — you've been on-call for high-traffic data systems, you know how to triage under pressure, and you write runbooks that actually get used

- Experience with container orchestration (Kubernetes) and deploying stateful workloads in production

- A bias for automation — if you've done something manual twice, you're already thinking about how to make it never happen again

- Strong written and oral communication skills, with the ability to translate infrastructure health into language product and business stakeholders understand

- The DNA to thrive in a fast-moving, high-autonomy environment — you see gaps as opportunities and own them end to end

Nice to Have

- Ownership of production database systems (Postgres, ClickHouse, Redis, or similar)

- Comfort reading and reasoning about Rust is a plus, as it's the language our database is written in

- Understanding of database reliability concepts — replication, backups, point-in-time recovery, connection pooling, and graceful degradation under load

Compensation

Salary Range: $180,000-$230,000 USD

Apply for this role →

← Back to all jobs