Principal Compiler Engineer
Job description
About the role
The is a critical role focused on the deep design and development of compiler infrastructure specifically tailored for demanding AI workloads. On this team, the successful hire will own the technical direction and implementation of compiler features that directly impact the performance and capability of the platform. The role requires close collaboration with cross-functional teams to translate high-level product requirements into robust compiler internals and optimization strategies. Technical judgment and risk assessments will be performed regularly to ensure that compiler decisions align with overarching product and business goals. The engineer will be responsible for ensuring a careful balance and priority across competing compiler features, performance targets, and development timelines. Day to day, the role involves hands-on design and development of compiler infrastructure for AI workloads to solve complex computational problems. The position demands a high level of ownership over the toolchain that powers AI inference and training efficiency on SambaNova hardware.
Key facts
What you'll do
You will architect the core components of the compiler stack to handle the unique dataflow patterns of AI models.
You will translate abstract product requirements into concrete compiler optimization passes and code generation strategies.
You will conduct rigorous technical judgment and risk assessments to evaluate the trade-offs of new compiler features.
You will ensure balance and priority across compiler features, performance targets, and resource constraints to meet release schedules.
You will debug and resolve complex compiler issues that arise from interactions with specific AI workloads and hardware capabilities.
You will collaborate with cross-functional teams on product requirements to define the compiler's behavior and interface contracts.
You will implement advanced optimization techniques focused on dataflow architectures to maximize throughput and efficiency.
You will mentor junior engineers on best practices for compiler development and systems programming within the codebase.
You will contribute to the selection and integration of compiler toolchains and third-party libraries relevant to AI processing.
You will analyze performance benchmarks to guide compiler decisions and validate the effectiveness of optimization strategies.
You will ensure that the compiler infrastructure supports the scalability and reliability required for production AI deployments.
You will participate in code reviews to maintain high standards of quality and maintainability in the compiler codebase.
You will work closely with hardware engineers to ensure that compiler output fully utilizes the capabilities of the target silicon.
You will document compiler internals and workflows to facilitate knowledge sharing and onboarding for new team members.
Requirements
- Bachelor's degree in Computer Science or related field, or equivalent experience.
- You need a minimum of 8 years of professional experience in compiler development, optimization, or related low-level programming.
- You need proven proficiency in systems programming languages, with C++ being a primary requirement.
- You need a deep understanding of AI/ML workloads, including common dataflow architectures and execution patterns.
- You need the ability to perform technical judgment and risk assessments for complex compiler design choices.
- You need experience with low-level code generation, instruction scheduling, and memory optimization techniques.
- You need strong debugging skills to diagnose and fix issues in compiler backends and optimization passes.
- You need the ability to work effectively in a fast-paced, iterative development environment with evolving requirements.
Nice to have
- Experience with ML frameworks and toolchains such as TensorFlow, PyTorch, or their derivative ecosystems.
- Knowledge of sovereign AI and regulated environments, including compliance considerations for data handling.
- Background in high-performance computing or distributed systems to understand scaling challenges.
- Familiarity with formal methods or program analysis techniques to improve compiler correctness.
- Experience with GPU or other accelerator programming models and their compilation challenges.
Skills & tools
You will use Compiler optimization techniques to generate efficient code.
You will work with Dataflow architectures to model and optimize computation graphs.
You will apply Systems programming skills to build and maintain the compiler infrastructure.
You will utilize static analysis tools to verify correctness and performance assumptions.
You will engage with version control systems to manage complex changes to the compiler codebase.
You will leverage profiling tools to identify bottlenecks in generated code and runtime behavior.
You will interface with hardware documentation to ensure compiler compatibility with target platforms.
Relevant systems
The work touches SambaCloud as the deployment target for compiled models and services.
The work touches SambaStack as the integrated software stack running on the compute platform.
The work touches SambaManaged for the managed services and orchestration layer.
The work touches SambaRack for the rack-scale infrastructure hosting the AI workloads.
The work touches RDU as a relevant unit or component within the system architecture.
The work touches Dataflow as a core architectural paradigm for AI computation.
The work touches Argyll as a specific component or codename within the stack.
The work touches Infercom as a tool or service related to compiler analysis or verification.
The work touches OVHcloud as a potential deployment or integration target.
The work touches SouthernCrossAI as a relevant project or initiative within the ecosystem.
Practical notes
This position requires work authorization in the United States.
The role is based in San Jose, California, and may require occasional travel to meet team obligations.
Equal Employment Opportunity; reasonable accommodations may apply during the hiring process.
The compensation details provided are indicative and subject to final negotiation based on experience and market conditions.
The role involves significant responsibility in shaping the compiler technology that drives AI infrastructure at scale.
Candidates must be able to demonstrate a track record of delivering complex compiler projects in a production environment.
The position requires a commitment to engineering excellence and collaboration within a distributed team structure.
The successful candidate will contribute to open source strategies where applicable and engage with the broader developer community.
This role reports to the Senior Director of Compiler Engineering or equivalent organizational leadership.
The position is expected to be filled as soon as possible given project timelines for AI platform development.
Interviews will likely involve technical assessments focused on compiler optimization and problem-solving skills.
The role requires a strong written and verbal communication skills to articulate technical trade-offs to both technical and non-technical stakeholders.
You will be expected to participate in on-call rotations for critical compiler issues impacting production deployments.
The role involves working closely with product managers to prioritize features based on market needs and technical feasibility.
You will need to manage multiple concurrent tasks and context switches without losing attention to detail or quality.
The role requires a proactive approach to identifying technical debt and proposing refactoring efforts to improve long-term maintainability.
You will be expected to write clean, modular, and testable code that adheres to the company's engineering standards.
The position offers opportunities for professional growth in the areas of compiler technology, AI infrastructure, and system-level programming.
You will engage with the latest research and developments in compiler theory and practice to apply them to real-world AI challenges.
The role is integral to the success of customers deploying agentic AI workloads in their own data centers using the SambaNova platform.