Platform Engineer
Your opportunity for impact:
NetDocuments is seeking an in-office LLM Platform Engineer who will replace our gateway with a real platform: a gateway that every internal model request flows through, with identity, budgets, routing, and audit built in.
You will own that gateway end to end — building it, running it, and being the person paged when it breaks. You will make AI spend attributable to the team and use case that generated it, and you will own the path a business team follows to get a new AI use case approved and provisioned. Our Corporate AI function owns AI strategy; you own the platform it depends on.
To be clear about where this sits: this is an internal business systems role. You will build and run the platform NetDocuments employees use to do their own work, not the customer facing product. It lives in Corporate IT Operations alongside identity, endpoint, and collaboration services, and it reports to the Senior Manager, Corporate IT Operations.
What your contributions will be:
AI Gateway Operations
Deploy, configure, and operate the AI gateway that fronts internal large language model traffic, including the provider connections behind it
Integrate the gateway with Microsoft Entra ID for single sign on and role based access, and manage virtual API keys through scoping, rotation, and revocation
Own gateway availability, latency, and capacity, including upgrades and change management
Implement request path controls such as redaction, data loss prevention, and audit logging to meet Security, Compliance, and FedRAMP requirements
Model Routing and Cost Control
Implement and tune routing policy so requests are served by the appropriate model by default rather than by user selection
Build spend attribution by team, application, use case, and API key across both seat based and API based consumption
Configure and enforce budgets, rate limits, and hard stops, replacing manually administered per person spend caps
Find and reduce avoidable consumption, including retry loops, oversized context, and wrong model defaults, and supply the usage data behind vendor and renewal decisions
Use Case Intake and Provisioning
Run the intake path that takes a team from an AI use case request to provisioned, governed access
Coordinate review with Security, Compliance, and Legal, and provision approved use cases with the right access, budget, and loggin
Maintain the register of approved use cases and drive down the time from request to access
Incident Response and Reliability
Act as first responder for gateway outages, provider degradation, authentication failures, rate limit events, and runaway spend
Maintain runbooks and escalation paths, and cross train the Corporate IT Operations team so gateway support is not single threaded
Run post incident review and drive corrective actions to closure
Reporting and Documentation
Provide a weekly AI platform update covering spend, adoption, gateway health, incidents, and risks
Produce monthly attribution and utilization reporting for Finance and leadership, and supply the underlying data to Business Intelligence for adoption dashboards
Document gateway configuration, provider connections, routing policy, and operational procedures
Other duties as assigned
Required experience and education:
Platform Operations & Reliability Engineering — 4-5 years operating a cloud production service that other teams depend on, including on call ownership and the habit of fixing causes rather than symptoms. AWS or Azure
API Gateway, Reverse Proxy & Service Mesh Engineering — hands on production experience with Envoy, Kong, LiteLLM, Portkey, NGINX, or similar, at the configuration and code level rather than through a vendor console
Authentication & Authorization Engineering — OIDC, OAuth 2.0, SAML, role-based access control design, and credential lifecycle management for both human and service identities
Backend & Systems Engineering — Python or Go at middleware level, plus infrastructure as code and containerized deployment
Observability & Instrumentation — metrics, structured logging, and distributed tracing, with the instinct to produce telemetry rather than only consume dashboards. Attribution is an instrumentation problem before it is a finance problem
Usage-Based Cost Attribution & Chargeback — you have made consumption spend attributable, forecastable, and enforceable. This role is measured substantially on financial outcomes
Process Design, Technical Documentation & Stakeholder Reporting — you automate coordination workflows rather than personally chasing them, and you write clearly for readers who are not engineers
Bachelor's degree — Computer Science, Engineering, or a related field, or equivalent practical experience
Ideally you will have:
Microsoft Entra ID specifically, as it is our identity provider; large language model API behavior including token accounting, streaming, and context limits; MCP and tool calling architectures; data loss prevention or data classification on a request path; SOC 2, ISO 27001, or FedRAMP environments; and SQL or Snowflake.
What You’ll Love About NetDocuments
The People!
90% healthcare premiums company covered
HSA company contribution
401K match at 4% with immediate vesting
Flexible PTO (typically 3 to 4 weeks a year)
10 paid holidays
Monthly contributions for life activities & wellness
Access to LinkedIn learning with monthly dedicated time to explore
Compensation Transparency
The compensation range for this position is: $110,000 - $120,000
The posted cash compensation for this position includes on target earnings, base salary and variable if applicable. Some roles may qualify for overtime pay. Individual compensation packages are determined based on various factors specific to each candidate, such as career level, skills, experience, geographic location, qualifications, and other job-related considerations
#LI-ONSITEEqual Opportunity