Solutions Architect
Job description
About the role
In this role you own the design and delivery of production-grade AI infrastructure and agent systems for LangChain's largest enterprise customers. You will translate ambiguous business problems into scalable, secure, and observable architectures that can be reliably operated in production. You own the full lifecycle of complex solutions, from initial discovery and technical assessment through design, implementation, validation, and ongoing optimization. You will work hand in hand with customers to align their strategic goals with the capabilities of LangChain's platform and open source ecosystem. You are expected to contribute directly to shaping best practices and reference architectures that influence product direction. This position demands comfort with ambiguity and the ability to dive deep into both infrastructure and agent-level implementation details. Your work will have direct impact on customer success and will help define how LangChain is used in production environments worldwide.
Key facts
What you'll do
You will design and deliver scalable, secure, and highly available infrastructure for AI platform deployments, including compute, storage, networking, and security controls. You will implement Infrastructure as Code using Terraform and Helm, and establish GitOps and CI/CD pipelines to automate deployment and operations. You will architect multi-region high-availability and disaster recovery strategies that meet strict enterprise requirements. You will design and implement multi-agent systems using LangChain and LangGraph, applying appropriate agent patterns to solve complex business problems. You will build comprehensive evaluation frameworks, conduct A/B testing on prompts and models, and optimize agent performance based on empirical results. You will integrate vector stores and RAG pipelines, manage stateful interactions, and implement tool integration with robust error handling. You will perform technical maturity assessments and infrastructure audits for enterprise customers, translating findings into clear recommendations. You will partner closely with Engagement Managers, Product, and Engineering to align solutions with product capabilities and roadmap. You will communicate technical trade-offs and recommendations to both technical and executive stakeholders. You will contribute to the evolution of LangChain best practices and reference architectures through hands-on customer work.
Requirements
7+ years of experience in technical, hands-on customer-facing roles such as Solutions Architect or Forward Deployed Engineer, and we also like former founders, so if you have an unusual background but all the right skillsets you are welcome to apply. 3+ years of experience designing and deploying production infrastructure on cloud platforms such as GCP, AWS, or Azure. Strong Kubernetes experience with GKE, EKS, or AKS, including cluster design, autoscaling, and multi-zone deployments. Experience with Infrastructure as Code tools including Terraform and Helm, and fluency in GitOps practices. Knowledge of database systems including relational databases and in-memory data stores, with attention to high availability, replication, backup strategies, and sizing. Experience designing high-availability and disaster recovery solutions that meet enterprise resilience standards. Strong understanding of networking, security controls such as SSO and RBAC, TLS, secrets management, and observability using tools like Prometheus, Grafana, and Datadog. Experience building and operating CI/CD pipelines for both infrastructure and applications. 1+ years of experience building production AI or ML applications and deploying agent-based systems. Strong hands-on experience with LLM frameworks such as LangChain and LangGraph for building agentic applications. Experience with state management patterns including both short-term and long-term memory. Experience designing and implementing evaluation frameworks tailored to AI applications. Strong prompt engineering skills with a track record of optimization and A/B testing in production. Experience with vector stores, RAG patterns, and knowledge organization strategies. Experience integrating tools, designing APIs, and implementing error handling patterns for agent workflows. Strong programming skills in Python and/or TypeScript required for implementing agents and infrastructure code. Customer-facing experience engaging enterprise customers and conducting technical assessments or infrastructure audits. Strong written and verbal communication skills to articulate complex technical concepts to diverse audiences.
Practical notes
Hours
Full-time
Travel
No travel is specified in SOURCE.
Visa
No specific visa information is provided in SOURCE.
Deadlines
No application deadline is specified in SOURCE.