Member of Technical Staff
Job description
About the role
SpaceXAI builds artificial intelligence systems designed to interpret the universe and support human knowledge. We maintain a flat, hands-on organization where every team member contributes directly to our core mission. This role is embedded within the RL infrastructure team, where your work will directly inform how we train and deploy intelligent systems. You will operate at the intersection of research and engineering, ensuring that the platforms we build can handle the demands of modern AI discovery. The focus is on creating robust, scalable foundations that allow our scientists to iterate quickly and reliably. You will be responsible for the end-to-end lifecycle of the systems you touch, from initial design through deployment and optimization. Your contributions will be visible and impactful, shaping the tools that push the boundaries of artificial intelligence. This is a position for someone who thrives on solving hard problems with precision and ownership.
Key facts
What you'll do
- Architect and implement the core infrastructure that powers reinforcement learning experiments from day one.
- Diagnose bottlenecks and failures in training pipelines, applying deep analytical skills to restore performance.
- Refactor legacy components to improve resilience, speed, and resource efficiency across the entire stack.
- Instrument systems with detailed telemetry to provide real-time insight into training jobs and infrastructure health.
- Collaborate closely with algorithm researchers to translate experimental requirements into stable, production-grade services.
- Evaluate and integrate new open-source tools and libraries to keep our technology stack at the forefront of the field.
- Optimize compute utilization and scheduling workflows to reduce waste and accelerate iteration cycles.
- Define and enforce best practices for configuration management and deployment automation.
- Mentor junior engineers by providing clear guidance and code reviews focused on quality and reliability.
- Partner with security and operations teams to ensure all deployments meet organizational standards and compliance requirements.
- Design fault-tolerant systems that can recover gracefully from hardware or software disruptions.
- Contribute to the documentation and design narratives that preserve institutional knowledge and accelerate onboarding.
- Lead incident response efforts during critical outages, driving toward permanent solutions rather than temporary fixes.
- Identify and eliminate technical debt to maintain a clean, sustainable codebase for long-term growth.
Requirements
- Demonstrated experience developing, debugging, and tuning distributed systems that operate at significant scale.
- Proven ability to navigate complex technical landscapes and resolve issues that span multiple layers of the software stack.
- Strong proficiency in Python, Jax, Rust, or C++, with a portfolio of work that showcases your skills in production environments.
- Comfort working in command-line environments and editing code directly on remote servers using standard tools.
- Experience with version control workflows using Git and a clear understanding of branching strategies and code review processes.
- Familiarity with Linux operating systems, including configuration, troubleshooting, and performance monitoring.
- A track record of taking ownership of projects and seeing them through from conception to completion without constant supervision.
- Willingness to adhere to strict engineering standards for code quality, testing, and documentation.
Nice to have
- Background in infrastructure designed for large-scale LLM training, including data loading, checkpointing, and networking patterns.
- Deep understanding of reinforcement learning methodologies, including exploration strategies, credit assignment, and policy optimization.
- Expertise in reinforcement learning numerics, such as precision choices, gradient scaling, and stability techniques.
Practical notes
The engagement is Full-time. Compensation ranges from $180,000 to $440,000 USD. The total rewards package includes equity, medical, dental, and vision insurance, 401(k) access, life insurance, and disability coverage. SpaceXAI is an equal opportunity employer.