Lead TechOps Engineer

Bybit · Kuala Lumpur, Malaysia · Engineering

Posted 2026-07-30

Apply for this role →

Job Responsibilities

Incident Response

Participate in 24/7 on-call rotation, responding promptly to major production incidents

Serve as the incident coordination hub, assembling relevant teams (R&D, SRE, Security, Risk, PR, CS, etc.), chairing emergency meetings, and driving business recovery

Determine incident severity and escalation timing, keeping management informed of incident impact, response progress, recovery expectations, and next steps

Unify the factual narrative when information is incomplete, and coordinate with PR and CS on external communications

Incident Post-Mortem & Risk Closure

Lead incident post-mortems: reconstruct timelines, document incident impact, response actions, and key decisions

Organize technical teams to complete root cause analysis, identify gaps in monitoring/alerting, system resilience, change management, and processes; drive incident classification, accountability assignment, and action item confirmation

Track action items to on-time completion, verifying effectiveness through testing, drills, or monitoring data

Regularly produce reports on incident trends, recurring issues, and major risks

Process & Capability Building

Establish and continuously optimize incident response SOPs, Runbooks, escalation paths, and compliance reporting processes; maintain on-call schedules and escalation chains

Organize emergency response training and cross-department drills; drive incident management tooling, data dashboards, and automation capabilities

──────

Who We're Looking For

Experience independently leading major incident response, including on-site coordination, management reporting, post-mortems, and risk closure

Ability to coordinate multiple technical and business teams without direct authority, continuously driving resolution under pressure

Strong technical comprehension — able to understand system architecture, service dependencies, monitoring/alerting, and incident chains, and judge whether root cause analysis and remediation plans are complete

Ability to quickly distill key information, deliver concise and clear briefings to management, take ownership of outcomes, and follow through until issues are closed

Strong written communication skills — able to produce well-structured post-mortems with factual evidence and clear conclusions

Experience with large-scale internet, fintech, payments, trading platforms, or other high-availability systems

Working proficiency in English — able to participate in English meetings and handle routine written communication

Willingness to accept 24/7 on-call rotation

Nice-to-Haves

Experience in crypto, exchanges, payments, or financial trading systems

Experience in SRE, production operations, technical support, or reliability engineering

Familiarity with PagerDuty, Datadog, Grafana, or similar monitoring and alerting tools

Familiarity with ITIL, Incident Management, or Problem Management frameworks

Apply for this role →

← Back to all jobs