Forward Deployed Engineer, LATAM (Remote)
The Role
We're looking for a Forward Deployed Engineer to join our LATAM team, with a focus on Bogotá, Mexico City, and São Paulo. In this remote, customer-facing role, you'll embed directly with enterprise customers across the region to architect and ship production systems on Telnyx's global network — voice, messaging, AI, and wireless — including self-hosted, open-weight LLM deployments that run inside customer environments. This isn't about demoing products; it's about building real solutions that work at scale.
This role would suit a technical professional who thrives at the intersection of engineering and customer success. You're equally comfortable debugging SIP traces as you are standing up a LiteLLM gateway with routing and failover across a customer's model providers, or leading a whiteboard session with a customer's engineering team. The ability to own outcomes from POC to production go-live is essential.
Customers across Latin America increasingly need Spanish- and Portuguese-first conversational AI, regional data residency, and flexible deployment options. A significant part of this job is making open-weight models perform reliably on infrastructure the customer controls. Professional fluency in Spanish or Portuguese, together with strong working proficiency in English, is required.
This is a LATAM-based remote role for candidates located in or near Bogotá, Mexico City, or São Paulo, with approximately 10–30% travel for deployments, workshops, and escalations across Colombia, Mexico, Brazil, and the wider region.
Responsibilities
Embed with enterprise customers to understand their communications workflows, AI use cases, and integration challenges firsthand
Build and deploy custom implementations: AI Voice Assistants, Telnyx APIs (Voice, Messaging, Fax, Wireless), and WebRTC
Deploy and operate open-weight LLMs such as Llama, Qwen, Mistral, DeepSeek, and gpt-oss in customer-controlled cloud, on-premises, and air-gapped environments
Deploy and operate LiteLLM as the model gateway in customer environments, providing a unified OpenAI-compatible interface across self-hosted models and hosted providers, with routing, load balancing, retries and fallbacks, rate limits, and per-team virtual keys
Instrument and govern LLM usage through cost tracking and budgets, caching, logging and observability using OpenTelemetry, Langfuse, or similar tools, and guardrails
Design model-routing strategies for real-time voice workloads, balancing latency, cost, reliability, and Spanish- or Portuguese-language quality
Make the build-vs-buy case between self-hosted open-weight models and hosted frontier APIs, while keeping customer application code portable across both
Adapt models to customer domains through prompt and RAG pipelines, LoRA/QLoRA fine-tuning, and evaluation harnesses for Spanish, Portuguese, and multilingual use cases
Lead POCs, pilots, and production launches from whiteboard to go-live
Own customer outcomes and remain engaged until the solution is live and stable
Collaborate directly with Product and Engineering to shape the roadmap based on field insights from the Latin American market
Create clear technical documentation, runbooks, and maintainable solutions for handoff in English and, where needed, Spanish or Portuguese
Troubleshoot and resolve complex integration issues alongside customer teams
Help customers address applicable privacy, security, telecommunications, and data-residency requirements, including Colombia's data-protection framework, Mexico's LFPDPPP, and Brazil's LGPD
What We Are Looking For
CS degree or equivalent experience
3+ years building software or doing technical consulting
Proficiency in multiple languages such as Python, Node.js, and Go — you're more dangerous in some than others
Hands-on experience running LiteLLM or a comparable LLM gateway such as Portkey, Kong AI Gateway, or an in-house proxy in production, including model definitions, proxy configuration, routing and fallback rules, and virtual key management
Practical understanding of production model-gateway challenges: provider rate limits and quotas, timeouts and retries, streaming, token accounting and cost attribution, and concurrency-related failure modes
Comfort deploying containerized services on Kubernetes, with secrets management, configuration, observability, and upgrades as part of the deployment story
Strong API fluency, event-driven thinking, and cloud-native instincts
Exposure to SIP, WebRTC, VoIP, or real-time voice and messaging systems
Ability to translate “it's not working” into root cause
Comfort working directly with customers and participating in high-stakes technical conversations
Professional fluency in Spanish or Portuguese
Professional working proficiency in English for internal collaboration, documentation, and work with Product and Engineering
Based in or near Bogotá, Mexico City, or São Paulo, with authorization to work in the applicable country
Willingness to travel approximately 10–30% across Latin America
Bonus Points For
Professional proficiency in both Spanish and Portuguese
Experience with AI voice assistants, STT/TTS, or LLM-based conversational systems, particularly for Latin American Spanish or Brazilian Portuguese
Fine-tuning and post-training experience: LoRA/QLoRA, distillation, preference tuning, or building evaluation sets for a specific domain
Experience with an inference-serving engine behind the gateway such as vLLM, SGLang, TGI, or Ollama, including sizing GPU capacity for self-hosted models
SQL proficiency with PostgreSQL, MySQL, or Oracle
ETL and data-wrangling experience
DevOps fundamentals including Docker, Kubernetes, and CI/CD
Background in telecom, CPaaS, contact centers, or high-growth SaaS
Experience designing private-cloud, sovereign-cloud, on-premises, or data-residency-sensitive architectures
Security mindset, including IAM, encryption, secrets management, and audit logging
#LI-RH1