Senior AI Platform Engineer ( Enterprise Systems)
About the Role
Senior engineer who ships production software and AI systems on cloud infrastructure that is reliable, secure, and cost-disciplined. You design applications and services, operate inference platforms and pipelines, and treat unit economics (tokens, GPU-hours, idle capacity) as engineering constraints.
You partner with product, engineering, and business stakeholders to deliver working software, higher reliability, and lower cost per unit of AI value.
What You Will Do
Software development
Design, build, and ship production services, APIs, and (where needed) user-facing interfaces.
Own the full lifecycle: implementation, tests, code review, CI/CD, deploy, and maintenance.
Write maintainable code in Python plus one more stack (TypeScript/Node, Go, or Java).
AI engineering
Build and operate production AI systems: RAG, evaluation, fine-tuning, serving, and inference optimization.
Turn prototypes into reusable platform capabilities with CI/CD, eval, and governance.
Cloud infrastructure
Architect AWS/GCP environments with Kubernetes, Terraform, and CI/CD for apps and model pipelines.
Own observability and incident response; partner on SOC 2/SOX and least-privilege IAM.
Cost control
Treat cloud and AI spend as an engineering problem: visibility, attribution, budgets, and reduction.
Optimize GPUs, inference routing, model choice, and right-sizing; track cost per request/token and GPU idle.
What We're Looking For
Required
5+ years of professional software development, owning production systems (not only scripts or notebooks).
Shipped backend services/APIs and production AI/ML (LLMs, RAG, serving).
Strong Python plus a second stack; PyTorch, TensorFlow, or equivalent serving.
AWS or GCP with Kubernetes, Terraform, and CI/CD.
Proven work reducing or governing cloud/AI spend.
Clear communication of technical and cost tradeoffs; you run what you build.
Preferred
Full-stack (React/Next.js), vector databases, inference optimization, FinOps tooling, SOX/SOC 2, Copilot/Cursor.
What Success Looks Like
Production software and AI with SLOs and real ownership — not demos.
Reliability stays high while GPU and cloud waste trends down.
Cost per unit of AI value is visible and improving; teams reuse libraries, IaC, and cost guardrails.