Forward Deployed Engineer, GCC
The Role
We're looking for a Forward Deployed Engineer to join our team in Riyadh or Dubai. In this role, you'll embed directly with enterprise and public-sector customers across the Kingdom to architect and ship production systems on Telnyx's global network — voice, messaging, AI, and wireless — including self-hosted, open-weight LLM deployments that run inside customer environments. This isn't about demoing products; it's about building real solutions that work at scale.
This role would suit a technical professional who thrives at the intersection of engineering and customer success. You're equally comfortable debugging SIP traces as you are standing up a LiteLLM gateway with routing and failover across a customer's model providers, or leading a whiteboard session with a customer's engineering team. The ability to own outcomes from POC to production go-live is essential.
Saudi customers increasingly require Arabic-first conversational AI and in-Kingdom data residency, so a large part of this job is making open-weight models perform in production on infrastructure the customer controls. This is a customer-facing role conducted largely in Arabic — fluent spoken Gulf Arabic is a hard requirement.
This is a Riyadh-based or Dubai-based role with 10–30% travel to customer sites for deployments, workshops, and escalations — primarily across the Kingdom and/or with occasional travel elsewhere in the GCC.
Responsibilities
Embed with enterprise and government customers to understand their communications workflows, AI use cases, and integration challenges firsthand
Build and deploy custom implementations: AI Voice Assistants, Telnyx APIs (Voice, Messaging, Fax, Wireless), WebRTC
Deploy and operate open-weight LLMs (Llama, Qwen, Mistral, DeepSeek, gpt-oss, and Arabic-first models such as ALLaM, Fanar, and Jais) in customer environments, including air-gapped and in-Kingdom sovereign cloud deployments
Deploy and operate LiteLLM as the model gateway in customer environments: a unified OpenAI-compatible interface across self-hosted open-weight models and hosted providers, with routing, load balancing, retries and fallbacks, rate limits, and per-team virtual keys
Instrument and govern LLM usage through the gateway — cost tracking and budgets, caching, logging and observability (OpenTelemetry, Langfuse, or similar), and guardrails — so customers can see and control what their AI workloads are doing
Design model routing strategies for real-time voice workloads, balancing latency, cost, and quality across Arabic-capable models, with sane fallback behavior when a provider degrades
Make the build-vs-buy case between self-hosted open-weight models and hosted frontier APIs, and keep the customer's application code portable across both
Adapt models to customer domains: prompt and RAG pipelines, LoRA/QLoRA fine-tuning, and evaluation harnesses for Arabic (including dialectal Arabic) and bilingual Arabic/English use cases
Lead POCs, pilots, and production launches from whiteboard to go-live
Own customer outcomes — stay engaged until the solution is live and stable
Collaborate directly with Product and Engineering to shape the roadmap based on field insights from the Saudi and wider MENA market
Create clear technical documentation, runbooks, and maintainable solutions for handoff, in English and where needed in Arabic
Troubleshoot and resolve complex integration issues alongside customer teams
Work with customers to meet local regulatory and data-residency requirements (CST/CITC, SDAIA, NCA, and PDPL obligations)
What We Are Looking For
CS degree or equivalent experience
3+ years building software or doing technical consulting
Proficiency in multiple languages: Python, Node.js, Go — you're more dangerous in some than others
Hands-on experience running LiteLLM (or a comparable LLM gateway such as Portkey, Kong AI Gateway, or an in-house proxy) in production — config-driven model definitions, the proxy server, routing and fallback rules, and virtual key management
Practical understanding of what breaks in front of a model in production: provider rate limits and quotas, timeout and retry behavior, streaming, token accounting and cost attribution, and the failure modes that only show up under concurrency
Comfortable deploying containerized services on Kubernetes, with secrets management, config, and upgrades as part of the deployment story
Strong API fluency, event-driven thinking, and cloud-native instincts
Exposure to SIP, WebRTC, or real-time voice/messaging systems
Ability to translate "it's not working" into root cause
Comfortable working on customer sites and in high-stakes technical conversations
Fluent spoken Gulf Arabic (Khaleeji) — required, not preferred. You'll run whiteboard sessions, live troubleshooting, and escalations in Arabic with customer engineering teams
Professional working proficiency in English for internal collaboration, documentation, and work with Product and Engineering
Based in Riyadh, or willing to relocate — this is a hybrid role with travel
Legally authorized to work in Saudi Arabia, or eligible for sponsorship
Bonus Points For
Experience with AI voice assistants, STT/TTS, or LLM-based conversational systems — especially Arabic ASR and speech synthesis
Fine-tuning and post-training experience: LoRA/QLoRA, distillation, preference tuning, or building eval sets for a specific domain
Familiarity with the Arabic open-weight model landscape and the trade-offs between Arabic-first and multilingual models
Experience with an inference serving engine behind the gateway (vLLM, SGLang, TGI, Ollama) and sizing GPU capacity for self-hosted models
SQL proficiency (Postgres, MySQL, Oracle)
ETL and data wrangling experience
DevOps fundamentals (Docker, Kubernetes, CI/CD)
Background in telecom, CPaaS, or high-growth SaaS
Experience with in-Kingdom sovereign or on-prem cloud deployments and PDPL-aligned architectures
Security mindset (IAM, encryption, audit logging)
#LI-RH1