Senior Software Engineer
Job description
About the role
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we are looking for passionate people to join our mission. If you are ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us. You will own the design and implementation of internal tools, specifically concentrating on Kubernetes Operators and Controllers to automate resource management and drive platform orchestration. You will lead complex, large-scale infrastructure shifts while taking ownership of incident response, authoring comprehensive retrospectives, and implementing systemic hardening to prevent recurrence using advanced overcommit strategies. In this capacity, you will operate as a "Captain," providing technical direction and unblocking team members across different time zones. Your work will ensure the reliability of our global fleet through a proven operations-heavy mindset focused on Day 1 and Day 2 activities for large-scale Kubernetes environments.
Key facts
What you'll do
Architectural Leadership by driving the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management and eliminate manual toil.
Platform Orchestration by leading complex, large-scale infrastructure shifts that modernize our global footprint and improve efficiency.
Operational Excellence through ownership of incident response, authoring comprehensive retrospectives, and implementing systemic hardening to prevent recurrence using advanced overcommit strategies.
Deepening Kubernetes Mastery by applying hands-on experience with custom operators and controllers that enforce desired state and automate cluster-level operations.
Advancing Tooling Proficiency to build production-grade tools and services that enhance infrastructure automation and deliver measurable reliability improvements.
Operating as a "Captain" in a multi-timezone environment, providing technical direction and mentoring junior engineers to maintain momentum and quality.
Leveraging Multi-cloud Infrastructure expertise to manage shifts across different environments, including familiarity with cloud-native tools and overseeing OS migrations.
Implementing Advanced Autoscaling and efficiency optimizations using tools like Karpenter to ensure cost-effective and resilient fleet management.
Utilizing Infrastructure as Code and Control Plane patterns with tools such as Terraform, Crossplane, and Cluster API to standardize deployments.
Employing GitOps and CI/CD practices by deploying and managing platform infrastructure via Argo CD to enforce declarative operations and traceability.
Requirements
5-8 years of experience in a DevOps, Site Reliability, or Infrastructure Engineering role that has prepared you for high-stakes platform ownership.
Deep internals knowledge of Kubernetes and hands-on experience writing custom operators that interact with the Kubernetes API server.
Strong experience building production-grade tools and services, specifically for infrastructure automation that scales with fleet size.
A proven track record of Day 1 and Day 2 operations for a large-scale Kubernetes fleet, including handling high-severity incidents and improving SLA compliance through automation.
Multi-cloud experience that includes familiarity with major cloud providers and managing OS migrations across diverse environments.
The ability to lead projects as a "Captain," providing technical direction and unblocking team members across different time zones to meet aggressive delivery timelines.
A commitment to fostering a diverse, welcoming, and inclusive environment where different backgrounds and abilities inspire better products and outcomes.
Willingness to comply with employment eligibility verification and security requirements, including criminal background checks and, when relevant, export compliance assessments.
Nice to have
A basic ability to read, navigate, and understand Golang code to collaborate effectively with engineering teams.
Experience implementing advanced autoscaling strategies and efficiency optimizations using tools like Karpenter.
Hands-on experience with Terraform or Crossplane and Cluster API for managing infrastructure as code and control planes.
Experience deploying and managing platform infrastructure via Argo CD for GitOps and CI/CD workflows.
Strong working knowledge of managing resources within AWS and Azure environments to optimize cost and performance.
Practical notes
Candidates must verify identity and eligibility to work and complete employment eligibility verification as part of the hiring process. A criminal background check is required. New Relic will consider qualified applicants with arrest and conviction records based on individual circumstances and in accordance with applicable law, including the San Francisco Fair Chance Ordinance. Employment eligibility verification and background checks are contingent on job offer. The role may require export compliance assessments for certain positions. This position adheres to New Relic's flexible workforce model, allowing for fully office-based, fully remote, or hybrid arrangements based on operational needs.