Solutions Architect
Job description
About the role
In this role you own the end to end design of production grade AI infrastructure and agentic systems for our largest enterprise customers. You translate ambiguous business problems into scalable secure architectures that leverage the LangChain platform and ecosystem. You own the full lifecycle of complex solutions from initial discovery through design implementation testing and ongoing optimization. You bridge the gap between cutting edge research and reliable production deployments by making pragmatic engineering tradeoffs. You also own the articulation of technical concepts to both technical and non technical stakeholders ensuring clarity and alignment. Your work directly influences how thousands of organizations build and operate intelligent agents in the real world.
Key facts
What you'll do
Architect and deliver scalable highly available infrastructure for AI platform workloads on GCP AWS and Azure including compute storage networking and security controls.
Design and implement multi agent systems using LangChain LangGraph and related frameworks applying robust patterns for orchestration routing and state management.
Lead technical discovery sessions and maturity assessments for enterprise customers extracting requirements and translating them into concrete technical roadmaps.
Construct comprehensive evaluation frameworks and A/B testing methodologies for LLM outputs ensuring reliability safety and performance metrics.
Implement Infrastructure as Code solutions using Terraform and Helm and establish GitOps workflows to automate deployment and operations at scale.
Optimize prompts and agent behaviors based on empirical data and experimentation results driving measurable improvements in accuracy efficiency and cost.
Integrate vector stores and implement RAG pipelines designing data ingestion chunking and retrieval strategies that support real world agent use cases.
Partner closely with Engagement Managers Product Managers and Engineering teams to align solutions with product vision and operational constraints.
Design and manage disaster recovery high availability and backup strategies for critical production agent workloads across multiple regions.
Implement monitoring logging and observability solutions using Prometheus Grafana and Datadog to ensure system reliability and rapid troubleshooting.
Guide customers on networking security and secrets management best practices including SSO RBAC TLS and secure storage patterns.
Collaborate with open source communities by contributing back to LangChain LangGraph and related projects shaping the broader ecosystem.
Develop and maintain detailed technical documentation runbooks and architecture diagrams to enable consistent delivery and knowledge transfer.
Support pre sales activities by building prototypes validating concepts and demonstrating the capabilities of the LangChain platform to prospective customers.
Requirements
7+ years of experience in technical hands on customer facing roles such as Solutions Architect or Forward Deployed Engineer with a track record of delivering complex solutions.
A background as a former founder is valued and if you have an unusual background but all the right skillsets you are welcome to apply.
3+ years of experience designing and deploying production infrastructure on major cloud platforms including GCP AWS or Azure with deep operational knowledge.
Strong expertise in Kubernetes cluster design and management on GKE EKS or AKS including autoscaling networking and multi zone high availability strategies.
Hands on experience with Infrastructure as Code tools such as Terraform and Helm and a proven history of implementing GitOps driven deployment pipelines.
Deep understanding of database systems including relational databases in memory data stores replication backup strategies and sizing for production workloads.
Demonstrated ability to design high availability disaster recovery and robust security architectures covering networking identity access control and secrets management.
Solid understanding of networking protocols security models including SSO RBAC TLS and secrets management practices as well as observability tools like Prometheus Grafana and Datadog.
Experience building production AI or ML applications and deploying them into operational environments with reliable performance and monitoring.
1+ years of experience building agentic applications using frameworks like LangChain LangGraph or similar libraries for stateful tool use and orchestration.
Experience with agent specific patterns including state management short term and long term memory tool integration and error handling strategies.
Proficiency in prompt engineering optimization and evaluation methods including A B testing and iterative improvement of agent behaviors.
Knowledge of vector stores RAG patterns and information retrieval strategies that support effective knowledge organization and retrieval.
Strong Python and or TypeScript development skills with the ability to write clean testable and maintainable code for complex systems.
Customer facing experience working directly with enterprise customers conducting technical assessments audits and presenting recommendations to leadership.
Exceptional written and verbal communication skills capable of translating technical concepts for both technical and non technical audiences.
Practical notes
The role is full time based in New York NY. Some travel may be required.
Candidates must be authorized to work in the United States without sponsorship now or in the future.