Lead TechOps Engineer
Job Responsibilities
Incident Response
Participate in 24/7 on-call rotation, responding promptly to major production incidents
Serve as the incident coordination hub, assembling relevant teams (R&D, SRE, Security, Risk, PR, CS, etc.), chairing emergency meetings, and driving business recovery
Determine incident severity and escalation timing, keeping management informed of incident impact, response progress, recovery expectations, and next steps
Unify the factual narrative when information is incomplete, and coordinate with PR and CS on external communications
Incident Post-Mortem & Risk Closure
Lead incident post-mortems: reconstruct timelines, document incident impact, response actions, and key decisions
Organize technical teams to complete root cause analysis, identify gaps in monitoring/alerting, system resilience, change management, and processes; drive incident classification, accountability assignment, and action item confirmation
Track action items to on-time completion, verifying effectiveness through testing, drills, or monitoring data
Regularly produce reports on incident trends, recurring issues, and major risks
Process & Capability Building
Establish and continuously optimize incident response SOPs, Runbooks, escalation paths, and compliance reporting processes; maintain on-call schedules and escalation chains
Organize emergency response training and cross-department drills; drive incident management tooling, data dashboards, and automation capabilities
──────
Who We're Looking For
Experience independently leading major incident response, including on-site coordination, management reporting, post-mortems, and risk closure
Ability to coordinate multiple technical and business teams without direct authority, continuously driving resolution under pressure
Strong technical comprehension — able to understand system architecture, service dependencies, monitoring/alerting, and incident chains, and judge whether root cause analysis and remediation plans are complete
Ability to quickly distill key information, deliver concise and clear briefings to management, take ownership of outcomes, and follow through until issues are closed
Strong written communication skills — able to produce well-structured post-mortems with factual evidence and clear conclusions
Experience with large-scale internet, fintech, payments, trading platforms, or other high-availability systems
Working proficiency in English — able to participate in English meetings and handle routine written communication
Willingness to accept 24/7 on-call rotation
Nice-to-Haves
Experience in crypto, exchanges, payments, or financial trading systems
Experience in SRE, production operations, technical support, or reliability engineering
Familiarity with PagerDuty, Datadog, Grafana, or similar monitoring and alerting tools
Familiarity with ITIL, Incident Management, or Problem Management frameworks