Sr. Site Reliability Engineer, Engineering Stack Support
About the role
Please note, this team is hiring across all levels and candidates are individually assessed and appropriately leveled based upon their skills and experience.
Engineering Stack Support (ESS) is a team of software engineers focused on improving availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning of the engineering stacks. In addition to that, ESS creates a bridge between service owners and platform engineering by applying a software engineering mindset to application debugging, system administration, and more.
What’s in it for you
You will be part of a growing team of renowned engineers in the exciting space of developing tools for infrastructure management.
Your contributions to our market-leading product support will significantly impact our rapidly-growing global customer base.
You will solve complex, exciting challenges and improve the depth and breadth of your technical and analytical skills.
What you will be doing
Partnering closely with our development teams and product managers to architect and build features that are highly available, performant and secure
Developing innovative ways to smartly measure, monitor & report application and infrastructure health
Gaining deep knowledge of our application stack
Improving the performance of micro-services and solve scaling/performance issues
Capacity management and planning
Wearing multiple hats in a fast-paced and rapidly-changing environment
Participating in 24X7 on-call rotations.
Required skills and experience
5+ years experience with troubleshooting Unix/Linux
Understanding of Networking concepts - TCP/IP, SSL/TLS, IPSec, GRE, VPN
Experience with algorithms, data structures, complexity analysis, and software design
Experience in one or more of the programming languages: Java, Python, Go
Experience in managing a large-scale web operations role
Hands-on experience with Ansible, Kubernetes, SQL and NoSQL datastores, CI/CD
Hands-on experience working with openstack , AWS or GCP cloud and automation
Experience with SumoLogic, Grafana/Prometheus and Elastic Stack.
Knowledge of distributed systems is a big plus.
Great written and verbal communication
Ability to work for a geo-distributed cross-functional group
Demonstrated ability to own and deliver projects independently
Strong interpersonal communication skills and the ability to work well in a diverse, team-focused environment with other SREs, developers, Product Managers, etc
Education
BSCS or equivalent required, MSCS or equivalent strongly preferred
#LI-JB3