Staff Software Engineer, Foundation Model API
Job description
About the role
Join the Foundation Model API team to develop the unified serving layer for large language models. You will work on infrastructure that supports both batch and real-time inference at an enterprise scale. In this role, you will own the design and execution of core components that enable seamless interaction between enterprise workloads and multiple model providers. You will be responsible for translating high-level product goals into robust, production-grade infrastructure that powers critical customer workflows. You will collaborate closely with cross-functional partners to ensure that the API layer meets stringent requirements for performance, reliability, and security. This position offers the opportunity to shape the foundational interfaces that data scientists and developers rely on when building AI on the Databricks platform. You will play a key role in defining how models are orchestrated, served, and monitored across distributed environments. Your work will directly influence the scalability and usability of AI services that support enterprise customers globally.
Key facts
What you'll do
- Engineer LLM infrastructure to support large-scale inference for partner models like Gemini, Anthropic, and OpenAI, as well as self-hosted options such as Llama, GPT-OSS, and Qwen.
- Define the product roadmap and execute strategy by engaging directly with enterprise users and model providers.
- Optimize distributed AI workloads for improved latency, efficiency, and reliability.
- Work alongside ML, platform, and infrastructure teams to create end-to-end user experiences.
- Influence how data scientists and developers interact with and build AI on the Databricks platform.
- Design and implement scalable APIs that abstract the complexity of model orchestration and deployment.
- Build observability and monitoring capabilities to ensure reliable operation of the model serving stack.
- Partner with security and compliance teams to enforce governance and access controls across AI services.
- Implement patterns for efficient resource utilization and cost management in AI workloads.
- Lead technical design discussions and contribute to architectural decisions that impact the entire platform.
- Collaborate with product managers to translate business requirements into technical specifications.
- Drive the delivery of high-quality features through code reviews, testing, and iterative improvements.
- Mentor engineers by sharing best practices and providing guidance on distributed systems and AI infrastructure.
- Contribute to open source initiatives and internal tools that enhance developer productivity and ecosystem integration.
Requirements
- 8+ years of professional experience in infrastructure or backend engineering.
- Proficiency in Python, Go, or Scala.
- Background in scalable APIs, distributed systems, or cloud-native infrastructure.
- Experience with GPU orchestration, ML infrastructure, or real-time model serving.
- Familiarity with system observability, deployment pipelines, and service-oriented architecture.
- Demonstrated ability to own products and ship value to end users.
- Strong understanding of networking, concurrency, and distributed computing principles.
- Experience working in agile environments and collaborating with cross-functional teams.
Nice to have
- Experience building products that support AI workflows.
- Exposure to cloud AI platforms such as Azure ML, Vertex AI, or SageMaker.
Practical notes
- Compensation includes potential equity, annual performance bonuses, and benefits.
- Databricks may choose not to proceed with applicants if the role requires access to export-controlled technology and a U.S. government license is not granted.
- This is a full-time position based in San Francisco, California.
- The role may involve collaboration with teams across different time zones.
- Applicants must be authorized to work in the United States without sponsorship for this position.
- No relocation assistance is provided for this role.
- The position requires availability during standard business hours for meetings and on-call responsibilities as needed.
- Travel is not expected for this role, but occasional in-person collaboration may be required.
- Visa sponsorship is not available for this position.
- Candidates must be able to start within a reasonable timeframe as defined by the hiring team.
- The application process may include technical assessments and interviews to evaluate core competencies.
- Interested candidates are encouraged to highlight relevant experience in AI infrastructure and distributed systems.
- All information provided in the job description is subject to change based on business needs.
- The successful candidate will be required to comply with Databricks security policies and procedures.
- This role is critical to the development of the Foundation Model API and will have direct impact on the company's AI product strategy.