DevOps / Site Reliability Engineer
Job description
About the role
General Matter is enriching uranium in America. Our mission is to restore our country's ability to make nuclear fuel. Our fuel will help power AI, manufacturing, and other critical industries. It will power our next generation of reactors. Ultimately, it will power our national ambitions. We were incubated by Founders Fund, like Anduril and Palantir before us, and we are backed by top tier investors. Our lean, world-class team of engineers and operators is applying a first-principles approach to solving the problem of nuclear fuel production. We are a mission-driven company with a culture of urgency, accountability and transparency. In this role, you own the design and operation of the software systems that ensure our enrichment infrastructure is observable, reliable, and efficient. You own the observability, alerting, and developer productivity platforms that keep our critical services online and our engineers moving fast. You own the incident response processes and long-term reliability initiatives that prevent failures before they happen. You will work closely with operators and engineers to instrument every service, automate critical workflows, and maintain the highest standards of production rigor. You are responsible for ensuring that both internal and production services are correctly monitored and that any issue is detected and resolved with speed and precision. Your work directly supports the safety and performance of the systems that enable our nuclear fuel production and national ambitions.
Key facts
What you'll do
Design, implement, and maintain observability and alerting systems across critical services and infrastructure.
Ensure all production and internal services are properly instrumented with metrics, logs, and traces to provide full visibility.
Own and maintain developer productivity tools, CI/CD systems, and internal platforms that accelerate safe delivery.
Participate in an on-call rotation and respond to production incidents with urgency and discipline to minimize impact.
Lead incident reviews and drive long-term reliability improvements by turning operational learnings into durable fixes.
Automate operational workflows to reduce manual toil and improve system resilience across the full stack.
Collaborate closely with engineers and operators to define service level objectives and ensure they are measurable and enforced.
Implement robust monitoring dashboards that reflect the health and performance of enrichment and supporting infrastructure.
Maintain configuration and deployment pipelines to ensure consistency, security, and repeatability across environments.
Contribute to the evolution of internal standards for logging, tracing, and alerting across engineering and operations.
Evaluate and adopt new tools and technologies that improve reliability, scalability, and operational efficiency.
Partner with security and infrastructure teams to ensure that monitoring and access controls meet regulatory and safety requirements.
Document operational runbooks and incident response procedures to improve clarity and speed of future responses.
Continuously refine alerting rules to reduce noise and ensure that critical issues are surfaced immediately and accurately.
Requirements
Demonstrate strong fundamentals in web service development and distributed systems with a proven track record of building reliable software.
Show solid understanding of networking concepts, including DNS, TLS/certificate management, and HTTP protocols and their practical implications.
Bring experience operating and debugging production systems in live environments where uptime and correctness are essential.
Show familiarity with observability tools such as metrics, logging, and alerting platforms, and how they fit into incident response.
Prove ability to write clear, maintainable code and automation scripts that other engineers can rely on and extend.
Demonstrate ownership, attention to detail, and sound technical judgment when making decisions that affect system reliability.
Commit to the ability to work extended hours and weekends as necessary to support critical operational needs.
Uphold the highest standards of professionalism, transparency, and accountability in a mission-driven, high-stakes environment.
Nice to have
Hands-on experience with modern observability stacks such as Prometheus, Grafana, OpenTelemetry, and Datadog.
Practical experience with cloud infrastructure and infrastructure-as-code tools to manage environments safely and at scale.
Exposure to CI/CD pipelines and developer tooling at scale, with a focus on automation and reliability.
Experience supporting safety-critical or high-reliability systems where failures have significant consequences.
Strong debugging skills across application, operating system, and network boundaries to isolate and resolve complex issues.
Prior on-call experience in a production environment, demonstrating comfort with responding to incidents at any time.
Practical notes
This is a full-time position based in Los Angeles, California.
The role may require extended hours and weekend availability as needed to support critical operations.
Employment is governed by merit, competence, and qualifications in accordance with our Equal Opportunity Employer policies.