Site Reliability Engineer
SpaceXUSA2w ago
Job description
About the role
In this role, you will be responsible for the development and upkeep of the infrastructure that supports the simulation and analysis of software for SpaceX's flight systems. Your work will encompass managing the entire lifecycle of high-performance web applications and distributed systems that are essential for the development of mission-critical hardware and software.
Key facts
What you'll do
- Automate the deployment processes and oversee the management of applications across both cloud-based and on-premises environments.
- Ensure the reliability and performance of core infrastructure components, including databases, messaging queues, storage solutions, and application servers.
- Optimize and scale RKE and Kubernetes clusters through the use of Ansible automation tools.
- Collaborate closely with software development and IT teams to design and implement maintainable, scalable products along with comprehensive test automation frameworks.
- Oversee the entire lifecycle of services, from initial design phases through to production operations.
- Collect and analyze requirements from engineering teams to facilitate the deployment of software platforms in a high-performance setting.
- Assume complete responsibility for the systems and tools you manage, contributing to the overarching goal of supporting multiplanetary exploration initiatives.
- Troubleshoot and resolve issues in production environments, ensuring minimal downtime and optimal performance.
- Develop monitoring solutions to track system performance and reliability metrics, enabling proactive maintenance.
- Document processes and create knowledge-sharing resources to enhance team collaboration and efficiency.
- Participate in on-call rotations to provide support for critical systems and respond to incidents as they arise.
- Engage in continuous learning and improvement practices to stay updated with the latest technologies and methodologies in site reliability engineering.
Requirements
- A Bachelor's degree in engineering, information systems, computer science, or a related field, or a minimum of 2 years of relevant professional experience in DevOps, software engineering, or site reliability engineering.
- At least 1 year of experience working with Linux operating systems, demonstrating a strong understanding of system administration.
- Proficiency in scripting languages such as Python or Bash, with the ability to write clean and efficient code.
- An active DOE Level Q clearance, or Top Secret/Top Secret SCI clearance is required.
- Candidates must comply with ITAR regulations, being a U.S. citizen, national, lawful permanent resident, asylee, or refugee.
- Strong problem-solving skills and the ability to work effectively under pressure in a fast-paced environment.
Nice to have
- Over 1 year of experience in Site Reliability Engineering, DevOps, or systems administration roles.
- Practical experience with containerization technologies, particularly Docker and Kubernetes.
- In-depth knowledge of Linux shell operations, including kernel modules, Control Groups, iptables, and Public Key Infrastructure (PKI).
- More than 1 year of experience in Python development, showcasing your ability to create robust applications.
- Familiarity with messaging systems like Kafka or RabbitMQ, enhancing your ability to manage data flows.
- Understanding of virtualization technologies, hypervisors, and techniques for optimizing database performance.
- Knowledge of identity management systems and authentication protocols, contributing to secure application environments.
- Experience with configuration management tools such as Helm, YAML, and Jinja, facilitating streamlined deployments.
Skills & tools
- Proficient in Linux, Kubernetes, Docker, Ansible, Python, Bash, RabbitMQ, Kafka, Helm, YAML, and Jinja, among other technologies.
Practical notes
- The salary range for this position is structured as follows: Level I ($125,000 - $145,000) or Level II ($145,000 - $175,000).
- Individuals holding active security clearances will receive a 10% pay differential (up to $20,000 annually) upon being briefed into classified programs.
- The total rewards package includes stock options, company shares, discretionary bonuses, and an Employee Stock Purchase Plan.
- Comprehensive benefits encompass medical, dental, and vision coverage, a 401(k) plan, life insurance, disability insurance, paid parental leave, three weeks of vacation, over ten holidays, and five days of sick leave annually.
- Be prepared for the possibility of working on weekends and extended hours as required by project demands.
- Positions that necessitate security clearance will be subject to pre-employment and random drug and alcohol testing to ensure compliance with safety regulations.