Staff Software Engineer - Infrastructure/DevOps
Role Summary
We are building a next-generation Core Infrastructure platform focused on:
Zero-trust security and identity-based access
Multi-region and multi-account scalability (multi cloud in future)
Highly automated, self-service infrastructure
Reliable and observable systems at scale
This role will own foundational infrastructure systems—networking, identity, compute platforms, and automation frameworks—that power all services.
Core Areas of Ownership
[1] Core Infrastructure & Cloud Platform
Design and evolve infrastructure on:
Amazon Web Services
Kubernetes
Build and scale:
Multi-region and multi-account architectures
Secure and scalable service connectivity (VPC, cross-account, cross-region)
[2] Security & Zero Trust
Drive adoption of:
Identity-based access (eliminate shared credentials)
Implement:
Secrets management using HashiCorp Vault
Policy enforcement using Open Policy Agent
Design:
Secure service-to-service communication
Fine-grained IAM and access control systems
[3] Infrastructure as Code & Automation
Build infrastructure using:
Terraform
Pulumi
Ansible
Create:
Reusable modules and abstractions
Fully automated provisioning workflows
Integrate with CI/CD:
GitHub Actions
Jenkins
[4] Cluster & Infrastructure Lifecycle
Own lifecycle of:
Kubernetes clusters (e.g., EKS)
Databases (eg, rds/aurora/elastic cache. etc)
Automate:
Cluster upgrades and migrations
Node scaling and patching
[5] Multi-Region & Migration Engineering
Design:
Region-agnostic infrastructure patterns
Enable:
Automated regions bootstrap and failover
Lead:
Migration initiatives (region, account, or cloud)
[6] Observability & Reliability
Build and operate:
Metrics, logging, and tracing systems using:
Prometheus
Grafana
ELK stack
Datadog
Improve:
System reliability, alerting, and incident response
[7 ] Governance & Standards
Define and enforce:
Resource tagging strategies
Data classification policies
Ensure:
Cost visibility and auditability
[8] Developer Productivity & Platform
Build:
Self-service infrastructure workflows
Improve:
Developer experience through automation and tooling
Contribute to:
Internal platform evolution
Responsibilities
Architect, build, and scale core infrastructure systems
Automate infrastructure lifecycle to minimize human intervention
Design secure connectivity across services, VPCs, accounts, and regions
Debug and resolve complex production issues across the stack
Write production-quality code, tools, and frameworks
Collaborate with engineering teams to standardize infrastructure usage
Minimum Qualifications
10+ years in a Software Engineering role or equivalent experience
5+ years of experience in infrastructure engineering
Strong hands-on experience with:
Amazon Web Services
Kubernetes
Terraform or Pulumi
Strong understanding of:
Networking (VPC, routing, DNS, connectivity)
IAM and access control
Experience with:
Automating infrastructure workflows
High-availability and scalable systems
Proficiency in:
Python / Go / Bash
Preferred Qualifications
Experience with:
Multi-region or multi-account architectures
Automating infrastructure migrations and upgrades (e.g., EKS upgrades)
Secrets management platforms (e.g., HashiCorp Vault)
Familiarity with:
Policy-as-code (e.g., Open Policy Agent)
Observability systems (Prometheus, Grafana, datadog)
Exposure to:
Data infrastructure (Hadoop, Trino, Spark)
Service mesh (e.g., Istio)