LANG CHAIN DEPLOYMENT AI ENGINEER ( LangGraph, LangSmith, Fleet, DeepAgents, Traces, Evaluations)

Nexaminds · India · Engineering

Posted 2026-08-21

Apply for this role →

About Nexaminds

Nexaminds is an AI and technology consulting company that helps organizations design, build, deploy, and operate production-grade generative and agentic AI solutions. We combine architecture, hands-on engineering, cloud delivery, and responsible AI practices to move customer initiatives from strategy and prototype to secure, production grade scalable business systems.

Position Summary

Nexaminds is seeking a hands-on LangChain Deployment AI Engineer to join our full-time consulting team and deliver customer-facing generative and agentic AI solutions. This role is designed for an engineer who can translate an approved agentic AI architecture into reliable production grade agents, work effectively alongside LangChain architects, and embed with customer engineering, platform, security, data, and product teams.

The engineer will build and deploy applications using LangChain and LangGraph; implement retrieval-augmented generation (RAG), agentic workflows, tool integrations, and evaluation pipelines; and own the operational details required to run those systems in production while monitoring agents using LangSmith. The successful candidate goes beyond demonstrations and notebooks. They can deliver tested services, secure integrations, automated releases, observability, cost controls, runbooks, and measurable production outcomes.

Role Mission

Convert LangChain solution architectures into maintainable, secure, observable, and scalable production deployments while accelerating customer adoption and transferring operational knowledge to customer teams.

Key Responsibilities

Customer deployment and delivery: Embed with customer project teams; clarify functional and nonfunctional requirements; convert architecture, user stories, and acceptance criteria into production-ready implementations; communicate progress, risks, tradeoffs, and dependencies clearly.

LangChain application engineering: Build LLM-powered assistants, copilots, chatbots, document-processing solutions, decision-support tools, and automated workflows using LangChain APIs, SDKs, and integrations.

Agentic workflow implementation: Develop deterministic and agentic workflows with LangGraph, including state management, routing, tool use, retries, checkpoints, human-in-the-loop controls, memory, streaming, and failure recovery.

RAG and enterprise data integration: Design and implement ingestion, chunking, metadata, embedding, indexing, retrieval, reranking, grounding, citation, and access-control patterns across structured and unstructured customer data.

Model and tool integration: Integrate hosted and open-source LLMs, embedding models, APIs, databases, search systems, vector stores, business applications, and custom tools using secure and maintainable interfaces.

Production deployment: Package services in containers; deploy to Kubernetes, serverless, or managed cloud runtimes; define configuration and secrets handling; support environment promotion, rollback, autoscaling, resiliency, and disaster-recovery requirements.

LangSmith observability and evaluation: Instrument traces, datasets, experiments, evaluators, feedback, alerts, and dashboards. Monitor model, tool, and retrieval behavior; latency; token usage; cost; error rates; and quality regressions.

LangChain enterprise delivery: Apply applicable LangChain enterprise capabilities, including LangSmith and Fleet, to support governed development, deployment, management, and operational visibility in customer environments.

Quality engineering: Create unit, integration, end-to-end, regression, adversarial, and load tests. Establish offline and online evaluation, golden datasets, quality thresholds, and release gates for prompts, models, retrieval, tools, and workflows.

Security and responsible AI: Implement least privilege, identity and access controls, encryption, secrets management, audit logging, data minimization, PII safeguards, prompt-injection defenses, output controls, content safety, and human escalation appropriate to the use case.

Performance and cost optimization: Improve response quality, latency, throughput, reliability, context utilization, caching, model selection, and token consumption while meeting agreed service-level objectives and budget constraints.

Operational readiness: Create deployment documentation, architecture updates, troubleshooting guides, runbooks, dashboards, alert thresholds, on-call handoffs, and incident/post-incident procedures.

Collaboration and knowledge transfer: Partner with LangChain architects, Nexaminds delivery leaders, and customer stakeholders; participate in design and code reviews; mentor engineers; and enable customer teams to operate and extend the delivered solution.

Typical Engagement Lifecycle

Discover and validate: Customer use case, environment and tech stack assessment, review the agentic AI architecture, use cases, data boundaries, security constraints, success metrics, dependencies, and deployment environment.

Build and integrate: Implement workflows, RAG, tools, APIs, persistence, identity, guardrails, and user-facing or service interfaces.

Evaluate and harden: Establish datasets and evaluators; test quality, safety, concurrency, failure modes, recovery, and cost.

Deploy and operate: Automate infrastructure and releases; configure observability, alerts, dashboards, runbooks, and service ownership.

Transfer and improve: Train customer teams, resolve early-life issues, analyze production feedback, and prioritize measurable improvements.

Required Qualifications

Bachelor’s degree in computer science, engineering, information systems, or a related discipline, or equivalent practical experience.

Strong professional software-engineering experience with Python and/or JavaScript/TypeScript, including API development, asynchronous processing, testing, packaging, dependency management, and code review.

Familiarity with the ADLC (Agent Development Life Cycle) and deploying configurable production-grade agents that can be managed and monitored in cloud platforms or on-prem datacenters.

Hands-on experience building agents with LangChain and LangGraph, including chains/runnables, tools, MCP servers, structured outputs, stateful workflows, streaming, persistence, and error handling.

Experience delivering LLM applications using commercial model APIs such as OpenAI, Anthropic, Google, or cloud-hosted equivalents, and/or open-source models.

Practical experience with embeddings, vector search, hybrid retrieval, metadata filtering, reranking, document ingestion, grounding, and citation patterns.

Experience with at least one vector database or search platform, such as Pinecone, Weaviate, Milvus, pgvector/PostgreSQL, Elasticsearch/OpenSearch, Redis, Azure AI Search, or a cloud-native equivalent.

Production experience with Docker and Kubernetes and with at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform.

Experience deploying production-grade agents and LLM applications beyond proofs of concept (POCs), including CI/CD, environment management, monitoring, incident troubleshooting, and operational support.

Hands-on experience with LangSmith or comparable LLM tracing, evaluation, and observability tooling; ability to investigate quality, latency, cost, and tool/retrieval failures from traces.

Working knowledge of application security, IAM, secrets management, encryption, API security, auditability, privacy, and secure software-development practices.

Strong customer-facing communication, technical writing, and consulting skills, including the ability to work through ambiguity and explain tradeoffs to engineers and business stakeholders.

Ability and willingness to travel to customer sites when an engagement requires it.

Preferred Qualifications

Experience with LangChain enterprise products and operating models, including LangSmith, Fleet and Deep Agents.

Experience designing or implementing multi-agent systems, supervisor patterns, long-running workflows, durable execution, event-driven integration, and human approval steps.

Experience with cloud AI platforms such as Amazon Bedrock, Azure AI Foundry/Azure OpenAI, Google Vertex AI, or managed Kubernetes platforms such as EKS, AKS, or GKE.

Experience with infrastructure as code and delivery tooling such as Terraform, Helm, GitHub Actions, GitLab CI, Jenkins, Argo CD, or equivalent platforms.

Experience with OpenTelemetry and observability platforms such as Datadog, Grafana, Prometheus, CloudWatch, Azure Monitor, or Google Cloud Operations.

Knowledge of LLM evaluation methods, model/prompt versioning, red teaming, adversarial testing, hallucination analysis, retrieval evaluation, and human feedback programs.

Experience in regulated or data-sensitive environments such as healthcare, financial services, insurance, manufacturing, or the public sector.

Relevant cloud, Kubernetes, security, data engineering, or AI/ML certifications.

Technical Competencies

Core frameworks: LangChain, LangGraph; LangSmith for tracing, evaluation, monitoring, and feedback workflows

Languages: Python; JavaScript and/or TypeScript; SQL; shell scripting

LLM engineering: Prompt and context engineering, structured output, tool/function calling, model routing, streaming, memory, caching, guardrails, evaluation

Data and retrieval: Document parsing, chunking, embeddings, vector/hybrid search, reranking, metadata and ACL filtering, relational and NoSQL data

Platform engineering: REST/GraphQL APIs, event queues, Docker, Kubernetes, CI/CD, infrastructure as code, secrets and configuration management

Operations: Distributed tracing, metrics, logs, dashboards, alerting, SLOs, load testing, incident response, cost management

Apply for this role →

← Back to all jobs