Sr. Kubernetes Platform Site Reliability Engineer
Job description
About the role
In this position, you will play a vital role in supporting the Starlink satellite network by designing and maintaining the infrastructure essential for delivering global broadband services. Your responsibilities will include managing extensive on-premise computing resources and collaborating closely with various engineering teams to ensure that our services remain highly available for millions of users every day.
Key facts
What you'll do
- Automate the deployment and management processes for on-premise Kubernetes clusters to enhance operational efficiency.
- Supervise core infrastructure elements, which encompass databases, monitoring systems, and distributed storage solutions.
- Collaborate with software engineers to create products that are both maintainable and scalable, ensuring they meet user demands.
- Oversee the complete service lifecycle, starting from the initial design phase through to operational management and ongoing enhancements.
- Set up monitoring and alerting systems to ensure optimal system availability and performance.
- Diagnose issues and integrate various components across the entire Starlink technology stack to maintain seamless operations.
- Identify and implement strategies to enhance system performance and reliability.
- Engage in proactive troubleshooting to resolve incidents and minimize downtime, ensuring a smooth user experience.
- Conduct regular assessments of system architecture and propose necessary upgrades or changes to improve efficiency.
- Mentor junior team members and share best practices in site reliability engineering and DevOps methodologies.
- Participate in on-call rotations to provide support during critical incidents and ensure rapid recovery from outages.
Requirements
- A bachelor's degree in computer science, information technology, or engineering, along with at least 5 years of professional experience in Site Reliability Engineering (SRE) or DevOps; or a minimum of 7 years of relevant experience without a degree.
- At least 2 years of hands-on experience with Linux operating systems, demonstrating a strong understanding of system administration.
- Proficiency in using infrastructure management tools such as Terraform or Ansible for automating infrastructure provisioning and management.
- Experience with containerization technologies, specifically OCI containers and Kubernetes, is essential.
- Strong scripting skills in languages such as Bash or Python, enabling automation of routine tasks.
- Background in software development with experience in languages like Python, C++, or Go.
- Willingness to work extended hours and weekends when necessary to support critical projects or incidents.
Nice to have
- Over 1 year of experience with Python-based development frameworks, enhancing your ability to contribute to software projects.
- Experience managing Kubernetes clusters at an advanced level, going beyond basic operational tasks.
- Familiarity with Linux boot processes and system configurations, contributing to a deeper understanding of system internals.
- Expertise in testing methodologies, continuous integration, deployment practices, and monitoring systems.
- Knowledge of build technologies such as Bazel or Makefiles, which can streamline the development process.
- Understanding of distributed databases and data modeling principles, aiding in effective data management strategies.
- A solid foundation in networking concepts, particularly TCP/IP, to troubleshoot connectivity issues.
- Strong communication skills to effectively interact with both technical and non-technical stakeholders.
Skills & tools
- Proficient in Kubernetes, Linux, Python, C++, Go, Bash, Terraform, Ansible, Bazel, Makefiles, and TCP/IP networking protocols.
Practical notes
- Salary range: $165,000.00 - $230,000.00, commensurate with experience and qualifications.
- Total compensation package includes company stock options, long-term cash incentives, discretionary bonuses, and an Employee Stock Purchase Plan.
- Comprehensive benefits package covering medical, vision, dental, 401(k) retirement plan, disability insurance, life insurance, paid parental leave, three weeks of vacation, and over ten paid holidays annually.
- Company-provided shuttle services are available from select locations in Seattle to the Redmond office, facilitating easy commuting.
- ITAR compliance is mandatory: Applicants must be U.S. citizens, lawful permanent residents, refugees, or asylees to be eligible for this position.
About the company
Based in Starbase, Texas, a company pursues work in spaceflight, telecom, and artificial intelligence. A unit handles rocket launches, leading yearly orbital counts beyond other national and commercial teams. Another unit runs Starlink, a satellite communications firm. A third area, through SpaceXAI, advances Grok and AI integration with X, while managing data center operations. Roles often cross these paths. Candidates with engineering, software, or operations backgrounds may find chances to join projects that span launch, connectivity, and machine learning efforts in this Texas-based enterprise.