Senior Software Engineer
Job description
About the role
You will architect the backend infrastructure that transforms how enterprise AI agents operate at scale within Snowflake's ecosystem. This role demands that you treat AI not as a feature but as a high-trust collaborator that reshapes execution strategy and product thinking. You will own critical pathways in agentic runtime orchestration, context retrieval, and evaluation frameworks that directly influence flagship offerings like Snowflake Intelligence. Your curiosity and low-ego approach will allow you to rapidly experiment, de-risk emerging patterns, and simplify complex workflows into robust production systems. You are expected to influence not just the code but the architecture that defines how agents behave in production. This position is for a builder who wants to leave a mark on the future of data-centric AI workloads. Your work will ensure that agentic workflows are fast, reliable, secure, and scalable for the world's largest enterprises.
Key facts
What you'll do
- Architect Agentic Runtimes: Build and scale the orchestration engines that execute complex agentic workflows, ensuring low-latency tool execution and robust state management across distributed environments.
- Scale Context Engineering Infra: Design high-performance systems for RAG (Retrieval-Augmented Generation), including vector database integration, scalable and efficient search indexing, query processing, result ranking, semantic caching, and automated metadata extraction for enterprise workloads.
- Build the "Evals Engine": Develop the automated infrastructure required to run massive-scale golden set simulations, error analysis pipelines, and "hillclimbing" experiments that continuously improve agent behavior and system reliability.
- Productionize AI Workflows: Collaborate closely with modeling teams to transform raw LLM capabilities into hardened, multi-tenant microservices that enforce strict guardrails, provide deep observability, and meet enterprise security standards.
- Optimize Performance & Cost: Direct the infra strategy for model routing, prompt caching, and token optimization to ensure Snowflake's AI features lead the industry in efficiency and cost-effectiveness for demanding workloads.
- Enable Cross-Team Experimentation: Partner with product and research teams to provision sandboxed environments where new agentic patterns can be prototyped, benchmarked, and validated before fleet-wide rollout.
- Drive Operational Excellence: Define and maintain service-level objectives, runbooks, and failure modes specific to agent execution paths, ensuring rapid detection, diagnosis, and remediation of issues in production.
- Champion Product Thinking in Infrastructure: Work alongside customer-facing teams to translate ambiguous problems into platform capabilities, distinguishing between one-off customer requests and infrastructure improvements that unlock value for the entire ecosystem.
- Implement Secure Multi-Tenancy: Build abstractions that isolate customer data and agent workflows, ensuring compliance and trust while maximizing resource utilization across shared cluster environments.
- Mentor and Influence: Act as a technical leader for junior engineers, shaping code review standards, design documents, and best practices for building reliable, scalable AI infrastructure.
Requirements
- Education: You hold a Bachelor's degree in Computer Science or a related technical field that provides foundational theory in algorithms, systems, and software design.
- Experience: You bring 7+ years of hands-on experience building distributed systems, high-throughput APIs, or backend infrastructure specifically for AI or ML products in production environments.
- Technical Stack: You demonstrate deep proficiency in Go or Java for systems-level work, and Python for AI orchestration, scripting, and integration with machine learning pipelines.
- Systems Thinking: You possess a strong understanding of database internals, distributed state management, consensus, and cloud-native architecture patterns including Kubernetes and systems like FoundationDB when applicable.
- Domain Expertise: You are fluent in the "plumbing" of AI, including vector indices, agent platforms, tool integration patterns, and building scalable data pipelines that move efficiently from ingestion to inference.
- Customer-Facing Experience: You have a track record of explaining complex technical failures to frustrated external audiences, earning their trust, and producing clear customer-safe summaries alongside internal analysis.
- Product Judgment: You can consistently determine whether a customer problem is a unique edge case or a platform gap that should be solved generically, and you are comfortable making that call with limited ambiguity.
- Cross-Layer Debugging: You excel at tracing a request across multiple services, correlating telemetry and logs, and root-causing issues in unfamiliar code without relying on guesswork.
- Eval Frameworks: You have experience defining quality metrics for LLM and agent systems, building eval harnesses, and using results to drive systematic improvements in model and infrastructure performance.
- Comfort with Ambiguity: You thrive in open-ended, externally-driven problem spaces where requirements evolve and success depends on experimentation and rapid iteration.
Nice to have
- Query optimization and deep knowledge of SQL engine internals that influence how analytical workloads are executed at scale.
- Experience designing multi-tenant systems that securely handle sensitive enterprise data while maximizing throughput and isolation.
- Background developing search infrastructure for large-scale applications, including relevance tuning, index lifecycle management, and semantic ranking.
- Direct, hands-on experience with any of the subsystems explicitly described in the responsibilities, such as agent runtimes, context retrieval pipelines, or eval execution frameworks.
Practical notes
This role is full-time and based in Menlo Park, California, United States. Compensation details are managed on the official Snowflake careers site and are not included in this description. Applicants must meet the stated eligibility requirements and are evaluated strictly against the outlined qualifications. The hiring process may involve multiple interview rounds assessing technical depth, system design, and product judgment. Reasonable accommodations may be available for individuals with disabilities during the recruitment process.