Mid-Level Site Reliability Engineer (SRE)
ROLE AT A GLANCE
We are looking for a Mid-Level Site Reliability Engineer to strengthen our platform and operations team. This role balances production reliability (incidents, SLOs, observability, on-call) with platform engineering (AWS infrastructure, Terraform, CI/CD, and container operations on ECS). The ideal candidate works autonomously on scoped services, collaborates closely with Engineering and QA, and contributes to continuous improvement of our operational practices.
Role: SRE / Platform OperationsLevel: Mid+Start: ASAPLanguages: ES + EN
Responsibilities
Reliability & Service Health: Support the definition and monitoring of service-level objectives (SLOs) and indicators (SLIs), working with Engineering teams to maintain production stability from deployment through ongoing operations.
Incident Response & On-Call: Participate in on-call rotations to triage, mitigate, and resolve production incidents; contribute to clear communication during outages and to blameless post-incident reviews.
Infrastructure as Code: Design, implement, and maintain AWS infrastructure using Terraform, following review practices and reusable patterns across environments.
CI/CD & Release Engineering: Collaborate with development teams to build and maintain CI/CD pipelines (e.g., Jenkins or GitLab CI/CD), enabling safe, repeatable deployments.
Observability & Monitoring: Configure and maintain monitoring and alerting with Datadog and AWS CloudWatch; build dashboards and tune alerts to reduce noise while preserving signal.
Container Operations: Operate and troubleshoot containerized workloads on Docker and Amazon ECS, including deployment, scaling, and runtime issues.
Automation & Operational Excellence: Develop scripts and tooling (e.g., Bash, Python) to automate repetitive operational tasks, reduce toil, and drive continuous improvement of platform processes.
Cross-functional Collaboration: Work closely with Engineering, QA, Product, and Security teams to align on requirements, change management, and improvements across the software development lifecycle (SDLC).
What You'll Bring (Requirements)
3+ years of professional experience in DevOps, SRE, or platform engineering roles, preferably within agile environments (Scrum, Kanban).
Cloud (AWS): Hands-on experience operating production workloads on AWS, including EC2, S3, RDS, IAM, and Cloud Watch.
Infrastructure as Code: Solid experience with Terraform; familiarity with AWS CloudFormation is a plus.
CI/CD: Experience building and maintaining pipelines with Jenkins or GitLab CI/CD;
strong Git workflow practices.
Containers: Practical experience with Docker and operating services on Amazon ECS.
Observability & Operations: Experience with Datadog and AWS CloudWatch; strong troubleshooting mindset for production incidents and performance issues.
Soft Skills: Proactive, adaptable, and collaborative mindset with the ability to work autonomously and effectively in multidisciplinary teams under time pressure.
Communication: Conversational Spanish and Intermediate English speaking skills are required
Seniority calibration: Ability to work independently on infrastructure changes and medium-severity incident response; ownership of scoped services and runbooks; active participation in post-incident reviews and platform improvements
Bonus Points (Desirable Experience)
Advanced Platform: Experience with Kubernetes or deeper AWS services (e.g., Lambda, CloudFront); exposure to AWS CloudFormation in production.
Observability Depth: Advanced Datadog usage, SLO-based alerting, or experience tuning observability for large-scale systems.
Mentorship: Willingness to support and collaborate with junior engineers through knowledge sharing, runbooks, and day-to-day
Benefits we offer
Access to e-Learning platforms.
Amazing people-oriented organizational culture
Working from anywhere
Challenging projects using the latest technologies with clients from the US.