Senior Software Engineer II, DevEx, OPX
SamsaraRemote (Canada; Remote - US)4w ago
Job description
About the role
Become a vital member of Samsara's Operational Excellence (OPX) team, which operates under the Developer Experience (DevEx) division. In this position, you will play a key role in enhancing the engineering landscape that supports our worldwide teams, with a strong emphasis on ensuring production health and reliability at scale. Your contributions will be essential in developing platform capabilities, observability tools, and automated safeguards that facilitate confident feature delivery and swift incident management.
Key facts
What you'll do
- Design and implement automated systems focused on production reliability and self-healing mechanisms, such as automated rollbacks and deployment safeguards, which will serve as essential tools for other engineering teams.
- Improve incident management tools and the experience of on-call engineers by minimizing alert noise and offering clear, actionable insights to enhance response times.
- Develop and optimize observability infrastructure, including monitoring, alerting, and performance tracking systems, to provide real-time visibility into system health.
- Contribute to the creation of AI-driven operational tools capable of detecting, addressing, and self-recovering from issues with minimal human oversight.
- Actively identify and eliminate operational inefficiencies to prevent incidents and enhance the on-call experience for engineers.
- Collaborate closely with product engineering teams to tackle reliability issues and advocate for best practices in service operations.
- Establish and champion operational excellence standards throughout the engineering organization, ensuring high-quality practices are maintained.
- Embody and promote Samsara's core cultural values within the engineering team, fostering a collaborative and innovative environment.
Requirements
- At least 8 years of experience in software product design and development, demonstrating a robust understanding of the software lifecycle.
- A Bachelor's Degree in Computer Science, Engineering, or a related discipline, or equivalent hands-on experience in the field.
- Minimum of 3 years of experience in infrastructure or platform engineering roles, showcasing a strong background in this area.
- Proven expertise in observability, reliability, operational metrics, and data analysis, with a track record of successful implementations.
- Experience in architecting monitoring systems, SLO platforms, and automated response workflows, utilizing tools such as Datadog or similar technologies.
- Background in building and maintaining large-scale enterprise software applications, ensuring high performance and reliability.
- Familiarity with Developer Experience (DevEx) principles and the development of internal tools that enhance engineering operations.
- Proficiency with cloud platforms like AWS or GCP, demonstrating an understanding of cloud infrastructure.
- Experience in applying AI-driven automation throughout the software development lifecycle, with a focus on improving task efficiency and delivery speed.
- Strong coding skills in languages such as Go or Python, particularly in relation to infrastructure, deployment, and operational challenges.
- Experience in providing technical mentorship and guidance to engineers, exemplifying best engineering practices.
- A proactive mindset focused on continuous improvement and a commitment to enhancing existing processes and workflows.
Nice to have
- Excellent collaboration and communication skills, enabling effective interaction across teams.
- Familiarity with incident management platforms such as Incident.io or PagerDuty, enhancing incident response capabilities.
- Experience with Infrastructure as Code (IaC) tools, especially Terraform, to streamline infrastructure management.
Skills & tools
- Proficient in monitoring and observability tools like Datadog (or alternatives such as New Relic, Grafana)
- Experienced with cloud services including AWS and GCP
- Skilled in programming languages such as Go and Python
- Knowledgeable in Infrastructure as Code (IaC) tools, particularly Terraform
- Familiar with incident management solutions like Incident.io and PagerDuty
Practical notes
- Annual Base
Salary: $154,700 - $260,000 USD
- Eligible for an initial RSU grant without a vesting cliff, along with ongoing refresh opportunities.
- Flexible remote work model, driven by employee preferences.
- Professional development stipend available to support continuous learning.
- Comprehensive health benefits and parental leave policies in place.
- Employment offers are contingent upon the candidate's ability to secure and maintain the legal right to work in the specified location.
- Please note that relocation assistance is not available for this position.
- The final stage of the interview process may require a full-day, in-person interview at the San Francisco office.