Cloud Operations Engineer
Job description
About the role
Join the global team maintaining the operational reliability of MongoDB Atlas, our multi-cloud database service. You will work within a 24/7/365 environment to support customers ranging from startups to large enterprises. In this capacity, you will own the design and execution of operational initiatives that ensure high availability and performance. The role requires a proactive mindset focused on refining how the team responds to and prevents service disruptions. You will be responsible for scaling internal operations through the implementation of robust tools and streamlined processes. A key part of the position involves driving down incident resolution times by designing and deploying improved systems and runbooks. You will analyze complex issues to determine root causes and translate findings into actionable improvements for the team. Collaboration with product and engineering partners will be essential to enhance the management applications that support the Atlas platform.
Key facts
What you'll do
- Coordinate with international colleagues to maintain uptime for the Atlas customer base by aligning on-call responsibilities and incident response playbooks.
- Scale operations by refining internal processes and implementing new tools that reduce manual effort and improve consistency.
- Design and deploy systems to decrease incident resolution times through automation, improved monitoring, and clearer escalation paths.
- Detect and resolve customer-facing incidents and automate routine troubleshooting to minimize service disruption and manual intervention.
- Perform root cause analysis after incidents to improve future workflows, ensuring that lessons learned are documented and acted upon.
- Document troubleshooting procedures and standard operating procedures to create clear guidance for current and future team members.
- Collaborate with product and engineering teams to enhance management applications, providing operational insights that drive better product decisions.
- Communicate status updates regarding major outages to executive leadership, ensuring transparency and clarity during critical events.
- Participate in a weekly on-call rotation to address proactive and reactive alerts, maintaining vigilance across all service regions.
- Monitor system health and performance metrics to identify trends, anomalies, and potential risks before they impact customers.
- Implement infrastructure as code practices to ensure that cloud environments are consistent, repeatable, and auditable.
- Optimize resource utilization and costs across cloud platforms while maintaining required levels of service reliability and performance.
- Serve as a technical liaison between operations and development teams to streamline the deployment and support of database services.
- Contribute to the development and maintenance of runbooks, ensuring they reflect the current state of systems and procedures.
Requirements
- Minimum 2 years of experience as an SRE, DevOps, or Cloud Operations engineer in a production environment.
- Proficiency in Linux system administration and troubleshooting, including command-line operations, process management, and file systems.
- Experience with system performance monitoring, data analysis, and reporting using tools and dashboards to inform decisions.
- Understanding of database operations and networking fundamentals like TCP/IP and DNS, including how they impact distributed systems.
- Familiarity with AWS, GCP, or Azure infrastructure, including core services used to deploy and manage cloud-native applications.
- Ability to write scripts or programs to address system issues, using languages such as Python, Bash, or similar automation tools.
- Proficiency in at least one of the following: Java, Go, or Javascript, including the ability to read and modify code for automation and integration tasks.
- A degree in Computer Science, Computer Engineering, or equivalent practical experience that demonstrates technical competence and problem-solving ability.
- Strong understanding of operational best practices, including incident management, change management, and service level objectives.
- Demonstrated ability to work effectively in a 24/7 on-call environment, managing alerts and responding to issues at any hour.
- Excellent written and verbal communication skills, with the ability to convey technical information to both technical and non-technical audiences.
- Commitment to following established processes while also identifying opportunities for improvement and innovation.
Nice to have
- Experience with MongoDB, Splunk, or Kubernetes.
Skills & tools
- Linux administration
- Networking (DNS, TCP/IP)
- Cloud infrastructure (AWS, GCP, Azure)
- Scripting (Java, Go, or Javascript)
- Monitoring and performance analysis
Practical notes
- This role requires occasional coverage outside of standard hours for regional holidays or offsites, which are scheduled in advance to ensure fairness and transparency.
- Benefits include competitive salary, equity, pension, health insurance, and 20 weeks of maternity/paternity leave.
- Regular performance and compensation reviews are provided to support professional growth and development.
- Req ID: 4263312827.