Site Reliability Engineer
Job description
About the role
SpaceX is on the lookout for a Site Reliability Engineer to join the Raptor engine program team. In this role, you will be responsible for managing essential infrastructure and tackling intricate systems challenges that aim to enhance the efficiency of rocket engine production. Your contributions will play a vital role in ensuring that our operations run smoothly and effectively.
Key facts
What you'll do
- Supervise the management of storage solutions, networking systems, high-speed interconnects such as Infiniband, and high-performance computing (HPC) environments.
- Take charge of the procurement, design, and integration processes for infrastructure hardware to ensure optimal performance.
- Work closely with propulsion engineers to identify and eliminate technical barriers that may hinder production.
- Collaborate with various IT and infrastructure teams across the organization to ensure cohesive operations.
- Enhance the performance of software utilized in manufacturing tools and applications, including StarCCM+ and ANSYS.
- Implement monitoring solutions to maintain system reliability and performance metrics.
- Develop and maintain documentation for systems and processes to ensure clarity and continuity.
- Participate in incident response and troubleshooting efforts to quickly resolve any issues that arise.
- Analyze system performance data to identify areas for improvement and implement necessary changes.
- Engage in continuous learning to stay updated on industry trends and best practices in site reliability engineering.
- Assist in the training and onboarding of new team members to foster a knowledgeable workforce.
- Contribute to the development of automation scripts to streamline operations and reduce manual intervention.
Requirements
- At least 1 year of practical experience working with enterprise networking, server and client hardware, virtualization, security, and management tools.
- A bachelor's degree in a relevant scientific discipline, mathematics, engineering, or computer science; or a minimum of 2 years of professional experience in software development.
- Strong familiarity with both Windows and Linux server environments, demonstrating versatility in system management.
- Proven ability to troubleshoot and resolve complex technical issues in a timely manner.
- Excellent communication skills to effectively collaborate with cross-functional teams.
- A proactive mindset with a focus on continuous improvement and innovation in processes.
Nice to have
- At least 1 year of experience in systems engineering, showcasing a solid understanding of system design principles.
- Proficiency in using automation tools such as Ansible or Puppet, along with scripting skills in Python or Bash.
- Experience in deploying and troubleshooting large-scale computing environments, demonstrating an ability to manage complexity.
- Familiarity with resource management technologies like Docker and Kubernetes to enhance operational efficiency.
- Background in finite element analysis (FEA) and computational fluid dynamics (CFD) applications.
- Capability to identify performance bottlenecks and design high-performance systems that meet demanding requirements.
Skills & tools
- Proficient in Linux, Windows Server, Bash, Python, Puppet, Ansible, Kubernetes, Docker, Infiniband, ANSYS, StarCCM+, high-performance computing (HPC), computational fluid dynamics (CFD), and finite element analysis (FEA).
Practical notes
- Compensation for this position varies based on experience and level, with Level 1 salaries ranging from $125,000 to $150,000 annually, while Level 2 salaries range from $145,000 to $175,000 annually.
- Comprehensive benefits package includes medical, dental, and vision insurance, a 401(k) plan, disability and life insurance, parental leave, three weeks of vacation, and over ten paid holidays.
- Total rewards may also encompass stock options, company shares, long-term cash awards, discretionary bonuses, and participation in an Employee Stock Purchase Plan.
- Please note that due to ITAR regulations, applicants must be U.S. citizens, lawful permanent residents, refugees, or asylees to be eligible for this position.
About the company
Based in Starbase, Texas, a company pursues work in spaceflight, telecom, and artificial intelligence. A unit handles rocket launches, leading yearly orbital counts beyond other national and commercial teams. Another unit runs Starlink, a satellite communications firm. A third area, through SpaceXAI, advances Grok and AI integration with X, while managing data center operations. Roles often cross these paths. Candidates with engineering, software, or operations backgrounds may find chances to join projects that span launch, connectivity, and machine learning efforts in this Texas-based enterprise.