Site Reliability Engineer (SRE)
Job description
About the role
You will own the design and operation of the core infrastructure that powers our LEO satellite network and ground control systems. This role requires you to build and maintain the observability, automation, and reliability frameworks that ensure our services remain available and performant. You will collaborate closely with software engineers, data scientists, and mission operations teams to translate business requirements into resilient technical solutions. You will be responsible for defining service level objectives, implementing incident response playbooks, and driving continuous improvement across all production systems. You will lead efforts to automate infrastructure provisioning, manage cloud resources, and optimize cost efficiency at scale. You will also contribute to security and compliance initiatives by implementing robust access controls and monitoring policies. In addition, you will mentor junior engineers and foster a culture of reliability and operational excellence across the organization. Your work will directly impact the ability of the company to deliver secure and actionable intelligence from space to global customers.
Key facts
What you'll do
Design and maintain the reliability, scalability, and automation of Espace's global satellite infrastructure and ground systems.
Implement observability strategies using metrics, logs, and traces to provide deep insight into system behavior and performance.
Collaborate with cross-functional teams to define and track service level indicators, service level objectives, and error budgets.
Automate the deployment, scaling, and operation of containerized applications and infrastructure components using infrastructure as code practices.
Partner with security engineers to ensure systems adhere to security policies, access controls, and data protection standards.
Lead incident management activities, including on-call responsibilities, post-incident reviews, and the continuous refinement of runbooks.
Evaluate and integrate new tools and technologies to improve the efficiency and resilience of the development lifecycle.
Optimize cloud resource utilization and infrastructure costs while maintaining high standards of performance and availability.
Develop and maintain monitoring dashboards and alerting systems to enable proactive detection and resolution of issues.
Support the growth of the engineering organization by documenting processes, sharing knowledge, and promoting best practices.
Work closely with mission operations teams to ensure that infrastructure components support real-time satellite operations and data downlink.
Contribute to the design of robust network architectures that can handle high-throughput data streams from low Earth orbit assets.
Drive automation for testing, validation, and rollout of infrastructure changes across development, staging, and production environments.
Act as a technical leader in reliability engineering, guiding architectural decisions and long-term platform strategy.
Requirements
Candidates must have a Bachelor's degree in Computer Science, Engineering, or a related technical field or equivalent practical experience.
You must have demonstrated experience with cloud platforms, infrastructure automation, and container orchestration tools.
You should possess strong scripting abilities in at least one high-level programming language such as Python or Go.
You must have a solid understanding of networking concepts, including TCP/IP, DNS, HTTP, and load balancing mechanisms.
You should have experience with monitoring and observability tools such as Prometheus, Grafana, or similar platforms.
You must be comfortable working in a fast-paced environment where priorities shift based on mission needs and technical challenges.
You should have a proven track record of collaborating with cross-functional teams in a mission-critical setting.
You must be willing to participate in on-call rotations and respond to production incidents in a timely manner.
Nice to have
Experience with satellite communications systems or ground station operations is highly valued.
Familiarity with LEO satellite constellations and space-based data services is a strong advantage.
Knowledge of security and compliance frameworks relevant to aerospace and telecommunications is preferred.
Experience with CI/CD pipelines, configuration management, and infrastructure provisioning tools.
Practical notes
This role is based in Saratoga, California and requires in-office presence during standard business hours.
Applicants must be authorized to work in the United States without sponsorship for this position at this time.
Travel is not required for this role under current business conditions.
This job posting is intended for candidates who are excited about building infrastructure at the intersection of space and terrestrial systems.