Service Delivery Manager
The role
We are looking for a Service Delivery Manager to support strategic hyperscale customers using bare metal and GPU cloud AI infrastructure in Nebius data centers.
This role acts as the technical and operational interface between the customer/platform and Nebius infrastructure teams, ensuring reliable service delivery, SLA compliance, and smooth operations of large-scale GPU clusters and bare metal environments.
You will coordinate across data center operations, network, hardware lifecycle, and infrastructure engineering teams to deliver world-class infrastructure services for large AI workloads.
Your responsibilities will include:
Customer Infrastructure Ownership
Serve as the primary technical point of contact for the customer.
Manage operational relationship with hyperscale customers.
Coordinate infrastructure lifecycle including provisioning, maintenance, and incident management.
Service Delivery & SLA Management
Ensure SLA and SLO compliance for infrastructure services.
Drive incident management and root cause analysis.
Track service performance and operational metrics.
Data Center & Infrastructure Coordination
Work closely with data center operations teams to manage hardware support and infrastructure maintenance.
Coordinate bare metal deployments, replacements, and capacity expansion.
Align infrastructure operations with customer workload requirements.
Operational Excellence
Establish operational processes for hyperscale infrastructure environments.
Lead service reviews and operational planning with internal and customer teams.
Improve reliability, response times, and operational workflows.
Cross-Functional Collaboration
Partner with engineering, networking, and hardware lifecycle teams.
Support infrastructure scaling and new cluster deployments.
Participate in planning for large-scale AI compute infrastructure.
We expect you to have:
5+ years in technical account management, service delivery, or infrastructure operations
Experience working with hyperscale customers or large enterprise clients
Background in cloud, AI infrastructure, or data center operations
Understanding of bare metal infrastructure
Experience with data center environments
Familiarity with GPU clusters / AI workloads (preferred)
Knowledge of networking, hardware lifecycle, and infrastructure monitoring
Strong stakeholder management
Ability to operate in high-scale infrastructure environments
Excellent communication between technical and business teams
It will be an added bonus of you have:
Experience working with hyperscale companies
Experience supporting AI / ML infrastructure
Experience with GPU clusters