Site Reliability Engineer, GNC
Job description
About the role
SpaceX is on a mission to enable human life on other planets, with a strong emphasis on creating reusable launch technologies and the Starlink satellite network. As a Site Reliability Engineer within the Guidance, Navigation, and Control (GNC) team, you will play a crucial role in overseeing and enhancing the essential software and hardware systems that are vital for vehicle design and flight operations. This position is integral to ensuring the reliability and efficiency of mission-critical systems.
Key facts
What you'll do
- Implement, manage, and optimize GNC software solutions and services to ensure high availability and performance.
- Oversee both virtual and physical server environments to maintain operational integrity.
- Collaborate with the High-Performance Computing (HPC) team to manage a cluster comprising tens of thousands of CPUs.
- Work closely with GNC software developers to create tools that are both maintainable and operationally efficient.
- Lead incident response efforts and monitor web services and applications to quickly address any issues that arise.
- Coordinate with IT to manage the GNC computational infrastructure effectively.
- Guide the entire service lifecycle, from the initial design phase through to deployment and ongoing maintenance.
- Provide expert technical support and troubleshooting for various GNC analysis applications.
- Set up automated deployment pipelines for web-based tools to streamline processes.
- Document software updates meticulously, including changes to operating systems and the rollout of new tools.
- Investigate and resolve performance bottlenecks to enhance system efficiency.
Requirements
- A bachelor's degree in computer science, information technology, engineering, mathematics, or a related scientific discipline, along with at least 2 years of experience in software development, OR a minimum of 4 years of professional experience in Site Reliability Engineering (SRE) or DevOps.
- At least 1 year of hands-on experience with Linux operating systems.
- Proficiency in Python and familiarity with Python-based frameworks for application development, with a minimum of 1 year of experience.
- Candidates must be U.S. citizens, lawful permanent residents, refugees, or asylees to comply with ITAR regulations.
- Willingness and ability to obtain a Top Secret security clearance as required.
Nice to have
- A background with 2 or more years in systems administration, SRE, or DevOps roles.
- Additional experience with Python and Linux for at least 2 years.
- Familiarity with containerization and orchestration tools such as Docker, Vagrant, and Kubernetes.
- Experience with configuration management tools like Ansible, Puppet, or Terraform.
- Knowledge of build systems such as Make, Bazel, Pants, Buck, or Gradle, as well as package managers like pip or npm.
- Understanding of virtualization technologies, hypervisors, databases, and data modeling concepts.
- Networking skills, particularly with TCP/IP protocols.
- Experience in scaling web applications and managing on-premises infrastructure, including GPU clusters.
- A background in high-performance computing or large-scale data analysis is advantageous.
Skills & tools
- Proficient in Linux, Python, Docker, Vagrant, Kubernetes, Ansible, Puppet, Terraform, Make, Bazel, Pants, Buck, Gradle, pip, npm, and TCP/IP networking.
Practical notes
- Salary for Site Reliability Engineer Level I ranges from $125,000 to $145,000, while Level II positions offer between $145,000 and $175,000.
- The total rewards package includes stock options, long-term cash incentives, discretionary bonuses, and participation in an Employee Stock Purchase Plan.
- Comprehensive benefits encompass medical, vision, and dental coverage, a 401(k) retirement plan, life and disability insurance, paid parental leave, three weeks of vacation, and over ten paid holidays annually.
- Candidates should be prepared for extended hours and weekend work as necessary to meet critical project deadlines.
- An active security clearance may be required, leading to involvement in sensitive missions, and candidates will be subject to pre-employment and random drug and alcohol testing.
About the company
Based in Starbase, Texas, a company pursues work in spaceflight, telecom, and artificial intelligence. A unit handles rocket launches, leading yearly orbital counts beyond other national and commercial teams. Another unit runs Starlink, a satellite communications firm. A third area, through SpaceXAI, advances Grok and AI integration with X, while managing data center operations. Roles often cross these paths. Candidates with engineering, software, or operations backgrounds may find chances to join projects that span launch, connectivity, and machine learning efforts in this Texas-based enterprise.