Cloud Site Reliability Engineer
Job description
Cloud Site Reliability Engineer at Bamboohr.
About the role
As a Cloud Site Reliability Engineer at Bamboohr, you will play a pivotal role in ensuring the reliability, performance, and scalability of our cloud-based systems. You will collaborate closely with cross-functional teams to enhance our infrastructure and streamline operational processes. Your expertise will be critical in driving initiatives that improve system availability and reduce downtime, ultimately contributing to the overall success of our cloud services. This position offers a unique opportunity to work in a dynamic environment where innovation and continuous improvement are highly valued.
Key facts
What you'll do
- Design, implement, and maintain robust cloud infrastructure solutions that meet the needs of our customers and support our services.
- Monitor system performance and reliability, utilizing advanced metrics and logging tools to identify and resolve issues proactively.
- Collaborate with software development teams to integrate reliability and performance best practices into the software development lifecycle.
- Automate operational processes to enhance efficiency and reduce manual intervention, leveraging tools such as Terraform, Ansible, or similar.
- Participate in on-call rotations to provide support for production systems, ensuring rapid response to incidents and minimizing downtime.
- Conduct root cause analysis for incidents and outages, developing strategies to prevent recurrence and improve system resilience.
- Develop and maintain documentation related to system architecture, operational procedures, and incident response protocols.
- Engage in capacity planning and performance tuning to ensure optimal resource utilization and system responsiveness.
- Stay current with industry trends and emerging technologies, recommending improvements and innovations to enhance our cloud offerings.
- Mentor junior engineers and contribute to a culture of knowledge sharing and continuous learning within the team.
- Collaborate with security teams to implement best practices for securing cloud environments and protecting sensitive data.
- Assist in the evaluation and selection of new tools and technologies to enhance our operational capabilities.
Requirements
- A minimum of 3 years of experience in a Site Reliability Engineering (SRE) or DevOps role, preferably in a cloud environment.
- Strong proficiency in cloud platforms such as AWS, Azure, or Google Cloud, with hands-on experience in deploying and managing cloud services.
- Solid understanding of containerization technologies, including Docker and Kubernetes, and experience with orchestration and management.
- Proficient in scripting and automation using languages such as Python, Bash, or Go to streamline operational tasks.
- Familiarity with monitoring and logging tools like Prometheus, Grafana, ELK Stack, or similar, with the ability to analyze and interpret data effectively.
- Experience with CI/CD pipelines and tools, such as Jenkins, GitLab CI, or CircleCI, to facilitate continuous integration and delivery.
- Strong problem-solving skills and the ability to work under pressure in a fast-paced environment.
- Excellent communication and collaboration skills, with a proven ability to work effectively in cross-functional teams.
- A degree in Computer Science, Engineering, or a related field is preferred, though equivalent experience will be considered.
Nice to have
- Experience with infrastructure as code (IaC) tools such as Terraform or CloudFormation.
- Knowledge of networking concepts and protocols, including DNS, TCP/IP, and load balancing.
- Familiarity with security best practices in cloud environments, including identity and access management (IAM) and data encryption.
- Previous experience in a startup or fast-growing company, demonstrating adaptability and a proactive mindset.
- Contributions to open-source projects or participation in tech communities related to cloud technologies.
Skills & tools
- Cloud platforms: AWS, Azure, Google Cloud
- Containerization: Docker, Kubernetes
- Automation: Terraform, Ansible, scripting languages (Python, Bash, Go)
- Monitoring: Prometheus, Grafana, ELK Stack
- CI/CD: Jenkins, GitLab CI, CircleCI
- Networking concepts and security best practices
Practical notes
- This position is based in Ottawa, Ontario, and may require occasional on-site presence.
- Candidates should be prepared to participate in an on-call rotation as part of the role.
- Bamboohr values diversity and inclusion and encourages applicants from all backgrounds to apply.
META
Company: Bamboohr
Title: Cloud Site Reliability Engineer
Listed
location: Ottawa, Ontario
Job type: Full-time
About the company
Healthcare in the U. S. is fundamentally broken.