Snr. Site Reliability Engineer (Remote in Brazil)
Please submit your resume in English.
To learn more about our team and office culture in São Paulo, Brazil, visit the following links.
Careers Page: https://www.knowbe4.com/careers/locations/sao-paulo
Glassdoor: https://www.glassdoor.com/Location/KnowBe4-S%C3%A3o-Paulo-Location-EI_IE969384.0,7_IL[…]M_-C1lsxoZq7Cx8IriVE8MkrzuTmnJzqego77RAWZz9sqGt_55BflwYKpQeg
LinkedIn: https://www.linkedin.com/company/knowbe4/life/brazil/
KnowBe4’s Site Reliability Engineers help ensure that our platforms are reliable, secure, scalable, and efficient. They work alongside other engineers in a fast-paced, agile development environment, and share solutions to advance the technologies running our systems, improve their safety and reliability, and make the complex distributed services that deliver our platforms easy to understand.
The ideal member of our team gets excited about new AWS service releases, stays up-to-date on industry trends and design patterns, and has excellent time-management and communication skills.
Some of the technologies we use:
Programming Languages - Python, Ruby, Rust
Infrastructure as Code - Terraform, AWS, OpenTofu
Source Code Management and CI/CD - GitLab, Git
Observability - DataDog
Containerized Workloads - Docker
Cloud-native infrastructure in AWS - ECS, Lambda, Step Functions, SNS/SQS, Transit Gateway, Aurora, DynamoDB, CloudFront, S3, AppSync, API Gateway, and many more.
Responsibilities:
Work with other Site Reliability Engineers to build highly scalable and resilient applications and infrastructure in AWS
Maintain and improve extensible infrastructure-as-code using Terraform
Learn, maintain, and improve our existing deployment strategies
Deliver effective observability, monitoring, and alerting patterns for KnowBe4’s applications and infrastructure
Act as an escalation point for identifying and resolving the root cause for production incidents
Provide assistance designing globally distributed systems and processes for the organization
Identify deficiencies in our current applications and infrastructure and correct them when found
Define new approaches and tailored solutions to complex technical problems
Act as a project leader with other Site Reliability Engineers and ensure progress is communicated effectively to project stakeholders
Minimum Qualifications:
BS/MS/Ph.D. or equivalent plus 5 years experience
Training in secure coding practices (preferred)
Proficient authoring scripts in one or more programming languages (e.g. Python, Ruby, Javascript).
Experience designing and operating high-scale patterns in AWS
Experience building and designing repeatable workflows for continuous integration and continuous deployment (CI/CD) - GitLab is preferred
Excellent communication skills
Effectively able to self-manage your time across competing projects
Ability to quickly understand and debug complex distributed systems
Additional Qualifications (Preferred):
Confident writing in Python
AWS Cloud Certification(s) - Professional Level
GCP and or Azure
Experience working for a public company
Open-source contributions or technical blog experience