Solutions Architect
Job description
About the role
At LangChain, the Solutions Architect operates at the intersection of infrastructure engineering and agentic AI design, owning the end to end architecture of production grade AI systems for enterprise clients. You will translate ambiguous business problems into concrete technical roadmaps that span cloud infrastructure, agent workflows, and evaluation methodologies. This role requires you to design, build, and validate scalable architectures while mentoring internal stakeholders on best practices for reliable agent deployment. You will directly influence how LangChain platforms are implemented in high impact, real world environments. Every decision you make shapes the reliability, security, and performance of customer agent systems.
Key facts
What you'll do
- Conduct in depth technical discovery sessions with enterprise customers to capture requirements, constraints, and success metrics for AI agent initiatives.
- Design scalable and secure infrastructure deployments using Infrastructure as Code patterns, defining modules for compute, storage, networking, and identity management.
- Architect multi region, highly available solutions that incorporate disaster recovery strategies, automated failover, and robust monitoring practices.
- Build and maintain evaluation frameworks that measure agent behavior, output quality, and business metrics, enabling data driven optimization.
- Implement vector store strategies and RAG pipelines that align with customer data governance, privacy, and latency requirements.
- Optimize language model prompts, tool use patterns, and agent orchestration logic through structured A B testing and iterative refinement.
- Partner with product, engineering, and engagement teams to define reference architectures that can be standardized across similar customer profiles.
- Translate complex technical concepts into clear narratives for both technical and executive audiences during workshops and design reviews.
- Own the deployment lifecycle for agent applications, including CI/CD pipelines, environment management, and release strategies.
- Troubleshoot production incidents by analyzing logs, traces, and agent execution paths to identify root causes and prevent recurrence.
- Guide customers on cost optimization, resource allocation, and scaling patterns to ensure sustainable operation of agent platforms.
- Document architectural decisions, integration patterns, and operational runbooks to ensure continuity and knowledge transfer.
Requirements
- Bring 7 or more years of hands on experience in technical, customer facing roles such as Solutions Architect or Forward Deployed Engineer, with a preference for those who have also operated as founders.
- Demonstrate at least 3 years of designing and deploying production infrastructure on major cloud platforms, including GCP, AWS, or Azure, with proven records of stability and performance.
- Show advanced competency in Kubernetes across managed services such as GKE, EKS, or AKS, including cluster design, autoscaling configurations, and multi zone high availability.
- Apply Infrastructure as Code methodologies using tools such as Terraform and Helm, supported by GitOps workflows for reliable and repeatable deployments.
- Manage database systems with experience in relational databases, in memory data stores, and associated strategies for high availability, replication, backup, and capacity planning.
- Design high availability and disaster recovery solutions that meet strict recovery objectives and resilience standards for enterprise workloads.
- Maintain strong fluency in networking concepts, security controls such as SSO, RBAC, TLS, and secrets management, as well as observability tools like Prometheus, Grafana, and Datadog.
- Build CI/CD pipelines that automate infrastructure provisioning, testing, and application deployment across environments.
- Accumulate 1 or more years of experience building production AI or ML applications, with a focus on agent based systems that leverage frameworks such as LangChain and LangGraph.
- Apply state management patterns effectively, including both short term and long term memory mechanisms within agent workflows.
- Construct evaluation frameworks that assess agent performance, correctness, and alignment with intended business outcomes.
- Exhibit strong prompt engineering capabilities, including optimization techniques, experimentation design, and measurement of agent behavior.
- Work with vector databases, RAG architectures, and knowledge organization strategies to ensure accurate and efficient information retrieval.
- Integrate third party tools and APIs into agent systems, handling errors, retries, and fallbacks in resilient ways.
- Write clean, tested code in Python and/or TypeScript, demonstrating solid software engineering practices in agent centric applications.
- Engage with enterprise customers through technical assessments, audits, and design discussions, articulating clear recommendations based on observed needs.
- Communicate in a clear, persuasive manner to diverse audiences, adapting style for technical engineers, leadership stakeholders, and business decision makers.
Nice to have
- Preferred experience with specific cloud provider certifications, observability platforms, or infrastructure automation tools that align with our current stack.
- Background contributing to open source projects related to agent frameworks, evaluation tooling, or infrastructure libraries.
- Familiarity with regulated industries such as finance, healthcare, or enterprise SaaS, where security and compliance considerations are paramount.
- Experience supporting customers with distributed teams across multiple time zones and coordinating asynchronous collaboration.
Practical notes
This role is full time based in Los Angeles, California. Applicants must be authorized to work in the United States without sponsorship now or in the future. Relocation is not available for this position. The position does not require travel.