Site Reliability Engineer
Overview:
The Site Reliability Engineer (SRE) is responsible for the reliability, performance, and scalability of Precisely's infrastructure platforms across CEDAR (CCX) — an on-premises, private cloud managed services environment; RapidCX (RCX) — an AWS cloud environment for SaaS-delivered customer communications management; and Hosted Managed Services (HMS) — an AWS cloud environment supporting managed client deployments.
This role bridges software engineering and systems operations, building automation, observability tooling, and reliability standards to ensure platform availability and operational excellence. SREs are enabling partners: they set reliability standards, define what 'reliable' looks like for each service, validate production readiness, and coach engineering teams on operational best practices. Engineering teams own the reliability outcomes of the services they build; the SRE ensures they have the standards, tooling, and guidance to meet them.
What you will do:
Set and maintain reliability standards across CCX, RCX, and HMS, including SLOs, SLIs, error budgets, alerting, logging, tracing, and MTTR improvements.
Build and maintain infrastructure-as-code, deployment automation, monitoring, and observability tooling using Terraform, Ansible, Datadog, Python, and Bash.
Guide CI/CD and deployment standards while partnering with engineering teams to embed reliability, scalability, backup, recovery, and failure-mode planning into service design.
Review designs and lead Operational Readiness Reviews to validate production and disaster recovery readiness.
Lead major incident response, including incident command and clear stakeholder communication.
Produce root cause analyses, identify recurring issues, and implement preventive automation to reduce manual effort and improve reliability.
Maintain runbooks, reliability backlogs, and shared knowledge on monitoring gaps, incidents, lessons learned, and operational risks.
Use Precisely-provided AI tools for infrastructure code, incident analysis, troubleshooting, testing, runbooks, and documentation.
Ensure infrastructure meets security, compliance, vulnerability-remediation, and data-protection requirements, including applicable SOC 2 and FedRAMP standards.
Coach engineers, contribute to cross-team reviews, stay current with SRE practices, and participate in the rotating on-call schedule for critical escalations and changes.
What we are looking for:
Required:
Educational requirements (equivalent work experience will be accepted in place of the education requirement): Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent practical experience.
Years of experience: 3+ years of systems or infrastructure engineering experience in an enterprise production environment.
Years of experience needed with specific skills: Not separately specified beyond the overall experience requirement above.
Specific technical or software skills required:
Strong proficiency with Linux (RHEL/Oracle Linux) in a multi-site, multi-environment context.
Hands-on experience with infrastructure-as-code tools: Terraform and/or Ansible.
Experience deploying and managing workloads in AWS (EC2, ECS, S3, VPC, IAM).
Proficiency with at least one scripting language (Python, Bash) for automation development.
Demonstrated experience building and maintaining monitoring and alerting systems (Datadog preferred).
Solid understanding of TCP/IP networking, DNS, load balancing, and distributed systems.
Experience with CI/CD pipeline design and deployment automation standards.
Strong analytical skills; ability to perform structured root cause analysis and post-incident review.
Ability to define SLOs and lead Operational Readiness Reviews (ORRs); comfortable partnering with engineering teams on production readiness.
Demonstrated ability to work cross-functionally with engineering teams on reliability standards and observability requirements.
Necessary certifications: None required (see Preferred Skills below).
Travel is required: No — approximately 0%.
AI Skills/Knowledge:
Active use of Precisely-provided AI tools (GitHub Copilot, Claude, or equivalent) for code development, troubleshooting, and documentation is a required baseline for this role, not a differentiator.
Apply AI tools for infrastructure-as-code development, incident analysis, and runbook authoring.
Use AI tools to accelerate troubleshooting and solution testing.
Maintain working fluency with Precisely-approved AI coding assistants as part of daily practice.
Preferred Skills (a plus but not required):
Experience with containerization and orchestration (Docker, ECS, Kubernetes).
Familiarity with GitOps workflows and source control best practices (Git, GitLab).
Knowledge of enterprise virtualization platforms in a hybrid cloud context.
Understanding of change management and ITIL operational practices.
Experience with enterprise security tooling (Qualys, CrowdStrike, Rapid7).
AWS Solutions Architect, SysOps Administrator, or DevOps Engineer certification.
#LI-KM1 #LI-Remote