Forward Deployment Engineer
Job description
About the role
The role owns the design, implementation, and operation of production generative AI applications directly within strategic enterprise customer environments using SambaNova's SN40L platform. You will architect and deploy LLM-powered workflows such as RAG, multi-agent systems, and fine-tuning pipelines that are tailored to specific customer data and business constraints. You own the full performance lifecycle of AI inference on SambaNova hardware, optimizing for throughput, latency, and accuracy while benchmarking against customer targets and competitor baselines. You act as the primary technical resolver, troubleshooting issues end-to-end across models, software stacks, and underlying hardware in production. You translate observed field behavior into clear product requirements and feedback for internal engineering and product teams. You collaborate with sales and solutions engineering teams to shape technical narratives, scope engagements, and run compelling demonstrations during proofs of concept. You create and maintain reusable artifacts such as accelerators, reference architectures, and playbooks that allow successful patterns to scale across multiple customer deployments. You represent SambaNova at customer executive briefings and industry events, presenting technical findings and influencing roadmap direction based on real-world evidence.
Key facts
What you'll do
- Embed directly with strategic enterprise customers to design, build, and deploy production GenAI applications on SambaNova's SN40L platform and SambaStack based product portfolio.
- Architect and implement LLM-powered workflows including RAG pipelines, multi-agent systems, fine-tuning workflows, and coding solutions tailored to each customer's data, infrastructure, and business goals.
- Optimize AI inference performance on SambaNova hardware by benchmarking model throughput, latency, and accuracy against customer requirements and competitor baselines.
- Troubleshoot and resolve production issues end-to-end across model, software, and hardware layers, acting as the first and last line of technical escalation in the field.
- Translate customer needs into clear product requirements and engineering feedback, serving as the primary voice of field reality to SambaNova's Product and Engineering teams.
- Partner with Account Executives and Solutions Engineers to shape technical sales strategy, scope engagements, and demonstrate platform differentiation during evaluations and proof-of-concepts.
- Develop reusable accelerators, reference architectures, and internal playbooks that scale learnings from one deployment to many.
- Present technical findings, architecture decisions, and roadmap input at customer executive briefings and internal forums, representing SambaNova at industry conferences and events.
Requirements
- 5+ years of hands-on engineering experience, with a strong record of shipping production AI or HPC software in enterprise or cloud environments.
- Demonstrated experience deploying and optimizing AI inference workloads on heterogeneous compute, including GPUs, DPUs, or similar accelerators.
- Strong proficiency in Python and at least one systems-level language such as C or C++ for performance-sensitive components.
- Solid understanding of Linux operating system internals, networking, and storage systems in distributed, high-throughput contexts.
- Hands-on experience with containerization and orchestration tools such as Docker and Kubernetes in production AI workloads.
- Practical knowledge of large language model architectures, inference optimization techniques, and deployment patterns for generative AI.
- Track record of working effectively with cross-functional stakeholders, including product management, sales, and executive leadership.
- Willingness to travel occasionally to customer sites and industry events as needed to support deployments and strategic engagements.
Nice to have
- Experience with SambaNova hardware, the SN40L architecture, or prior exposure to dataflow computing paradigms.
- Familiarity with open-source model fine-tuning frameworks and productionization toolchains.
- Background in building or operating cloud-agnostic or on-premises AI inference platforms.
- Participation in AI or HPC community projects, open-source contributions, or published work in related technical domains.
Practical notes
- Travel may be required occasionally to support customer deployments and industry events.
- Visa requirements may apply for non-Indian candidates depending on specific engagement terms.
- This role is based in Bengaluru, Karnataka, India, and is subject to the constraints and expectations outlined in the source description.