Site Reliability Engineer - HPC & Automation
SpaceXUSA1w ago
EngineeringReliabilityAutomationremotecurated-jd
Job description
Site Reliability Engineer - HPC & Automation at SpaceX.
About the role
Join the Silicon Engineering team to manage the high performance computing infrastructure that powers Starlink chip development. You will build and scale the automated systems that enable rapid simulation and design cycles for our satellite constellation hardware.
Key facts
What you'll do
- Maintain, scale, and upgrade the cluster environments and services.
- Develop turnkey automation for silicon simulation workflows to improve project speed.
- Manage infrastructure as code and implement observability tools to monitor system health.
- Oversee continuous integration pipelines, build systems, and version control.
- Analyze performance metrics to identify and resolve system bottlenecks.
Requirements
- Bachelor degree in computer science, information systems, or an engineering field OR 2+ years of professional experience in system administration, HPC, or SRE.
- 1+ years of experience with Linux.
- 1+ years of development experience using Python, Bash, or similar languages.
- Ability to work extended hours and weekends to hit project milestones.
- Must meet ITAR requirements: U.S. citizen, lawful permanent resident, refugee, or asylee.
Nice to have
- Experience with containerization like Docker or Kubernetes.
- Knowledge of computer architecture, operating systems, and concurrency.
- Database management skills with MySQL, PostgreSQL, or SQLite.
- Networking proficiency in TCP/IP.
- Familiarity with workload managers such as Slurm or LSF.
- Experience with automation frameworks like Terraform, Ansible, or Puppet.
- Monitoring and alerting experience using Prometheus or Grafana.
- Background in CI/CD tools like Jenkins or Bamboo.
- Experience with networked storage automation and REST APIs.
- Familiarity with ASIC design tools from Cadence, Synopsys, Ansys, Keysight, or Siemens.
- Interest in AI or LLM-assisted tooling such as Grok or Claude Code.
Skills & tools
- Linux, Python, Bash, HPC, CI/CD, Infrastructure as Code, Observability, Networking, ASIC design flows.
Practical notes
- Salary range: Level 1 ($125,000 - $150,000) or Level 2 ($145,000 - $175,000).
- Compensation includes potential stock, long-term cash awards, discretionary bonuses, and an Employee Stock Purchase Plan.
- Benefits include medical, vision, dental, 401(k), life insurance, disability insurance, paid parental leave, 3 weeks of vacation, and 10+ paid holidays.
- Company shuttles operate between select Seattle locations and the Redmond office.