Senior Software Engineer - CI/CD
What You'll do:
Design & Develop: Build and operate the self-hosted GitHub Actions runner fleet on GKE including autoscaling, reliability tuning, and zombie-runner cleanup.
GitOps & Delivery: Own the ArgoCD topology powering CI/CD deployments — central architecture, cluster connectivity, ApplicationSets, production reliability (PDBs, replicas, spot-node avoidance), and event-driven flows via Argo Events + GCP PubSub.
Shared Platform Components: Maintain and evolve the Helm charts and ApplicationSet patterns used by every Kubernetes workload — versioning, release process, backward compatibility, and developer ergonomics.
Jenkins Platform Ownership: Own and evolve our Jenkins environment including shared libraries, controllers and agents, plugins, credentials integration, JVM upgrades and platform reliability.
Self-Service Tooling & Migration: build resusable GitHub actions workflows, actions, and templates, while driving migration from Jenkins and partnering on shared developer tooling
End-to-End Ownership: Lead complex platform projects from architectural design through implementation and long-term maintenance — identifying system-wide bottlenecks before they become outages, and owning the outcomes well past the ship date.
Operational Excellence: Own pipeline reliability end-to-end — observability with Datadog (dashboards, monitors, runner log analysis, cluster tracing), incident response and on-call rotation via PagerDuty, plus secrets rotation and vulnerability response.
Global Collaboration: Partner with distributed teams of architects, infrastructure and security engineers, and every product team that consumes the platform — coordinating across time zones to translate their pipeline needs into golden paths, reusable workflows, and self-service tooling.
What you'll bring:
CI/CD Platforms: Deep expertise in Jenkins administration, Groovy shared-libraries, Docker-based agents, GitHub Actions, reusable workflows, self-hosted runners, secrets management, and pipeline troubleshooting.
Cloud & Containers: Strong Kubernetes(GKE) administration, cluster optimization, networking, kubernetes internals, and hands on experience with GCP and AWS services
GitOps & Delivery Tooling: Hands-on experience with ArgoCD (ApplicationSets, sync strategies, alerting), Helm chart authoring and versioning, and Argo Events or comparable event-driven delivery patterns.
Infrastructure as Code: Strong proficiency with Terraform for GKE clusters, GCP resources, and CI/CD-related modules.
Observability: Hands-on experience instrumenting and operating pipelines with Datadog — custom metrics, dashboards, monitors, and distributed tracing.
Incident Response: Experience running production on-call with PagerDuty — rotation management, alert routing, escalation policy design, and post-incident review.
Artifact & Image Management: Experience with container registries — JFrog Artifactory (including Xray/Curation scanning) and GCP Artifact Registry — and image lifecycle policies.
Programming: Comfortable shipping production code in at least one of Go, Python, or Node.js — enough to build controllers, GCP Cloud Functions, glue services, and AI-assisted pipeline tooling.
AI Development: Experience or knowledge building or deploying LLM-based applications, AI-assisted developer tooling, or managing AI infrastructure for engineering workflows.
#LI-Remote