Senior Customer AI Engineer - Token Factory
The role
We're looking for a Senior Customer AI Engineer who will join our Token Factory team to help our customers successfully transition from proof-of-concept to production and scale their AI workloads on Nebius infrastructure.
This role sits at the intersection of engineering, delivery, and customer success – ensuring that what was promised during pre-sales actually works reliably in production. You will work closely with customer engineering teams, Solution Architects, and Product/Infrastructure teams to drive stable, performant, and cost-efficient deployments.
This role is NOT:
A sales role (though you will support expansion through value)
A pure support role (you won’t just react to tickets)
A solution architect role (you won’t design systems from scratch)
You’re welcome to work remotely from Singapore.
Your responsibilities will include:
Own the production journey
• Lead the transition from PoC to production
• Ensure customer workloads are deployed, stable, and scalable
• Drive time-to-production and time-to-value
Ensure technical success in production
• Understand customer architectures and use cases
• Monitor and improve:
o performance (latency, throughput)
o cost efficiency
o reliability
• Identify and resolve bottlenecks proactively
Act as a trusted technical partner
• Work directly with customer engineering teams
• Provide guidance on best practices and optimisation
• Translate technical challenges into actionable solutions
Manage risks and incidents
• Act as a primary technical contact for production issues
• Coordinate with internal teams to resolve incidents
• Communicate clearly during high-pressure situations
Drive continuous improvement
• Identify opportunities to optimise and expand usage
• Provide structured feedback to Product and Infrastructure teams
• Help shape better solutions based on real customer needs
We expect you to have:
Technical background
Practical knowledge of inference frameworks (e.g. vLLM, TensorRT, or similar)
Solid understanding of:
- cloud or infrastructure systems
- distributed systems or high-load applications
- AI/ML workloads (LLMs, inference, etc.)
Ability to troubleshoot and reason about system performance
Customer-facing experience
Experience working directly with technical customers (e.g. engineers, ML teams)
Ability to communicate complex topics clearly and effectively
Ownership & execution
Strong sense of ownership – you drive outcomes, not just tasks
Ability to manage multiple customers and priorities
Structured, proactive, and solution-oriented mindset
It will be an added bonus if you have:
Experience with GPU workloads or AI infrastructure
Background in solutions engineering, SRE, or technical support in B2B environments