DevOps & Cloud Infrastructure Engineer
Role Overview:
We're seeking an experienced DevOps & Cloud Infrastructure Engineer to own and optimize our fully cloud-native AWS environment. This role is responsible for infrastructure-as-code, reliability, security and cost management across the platform, together with the identity, networking, database, mail and observability services that support it.
Key Responsibilities:
Own and operate the AWS estate — multi-account structure, ECS Fargate workloads, VPC networking, site-to-site VPN and Route 53 DNS.
Manage infrastructure-as-code in Terraform, including state ownership, module design, and bringing remaining manually created resources under code.
Administer cloud identity and access across AWS IAM, Cognito and Entra ID, including least-privilege role design and MFA coverage.
Operate Aurora PostgreSQL — backup policy, point-in-time recovery, restore drills, version upgrades and failover testing.
Run and extend the self-managed observability platform; define SLOs, alert routing and incident response.
Own cloud cost management (FinOps) — rightsizing, savings-plan coverage and spend visibility.
Manage AWS SES mail operations including DKIM/SPF/DMARC alignment, deliverability and bounce handling, plus Microsoft 365 tenant administration.
Maintain security posture — CVE remediation, penetration-test findings, secrets management and rotation, and GDPR/DSGVO compliance.
Build and maintain CI/CD pipelines in Azure Pipelines for container build and deployment.
Required Skills & Experience:
Deep hands-on AWS administration in production, including incident response and cost optimization.
Strong Terraform / infrastructure-as-code experience, including state ownership and module refactoring.
Solid networking and security fundamentals — VPC design, VPN, DNS, TLS certificates and firewall configuration.
Managed database operations experience (Aurora/RDS PostgreSQL or equivalent) — backups, restores and upgrades.
Container and Linux operations experience — Docker, ECS/Fargate or equivalent orchestration.
Experience with observability tooling (Grafana, Prometheus, Datadog) and SLO-based alerting.
Working knowledge of Azure, Microsoft 365 and Entra ID administration.
6–10 years in cloud operations, IT infrastructure or platform engineering.