Senior Site Reliability Engineer
Job description
Senior Site Reliability Engineer at Bamboohr.
About the role
As a Senior Site Reliability Engineer at Bamboohr, you will play a pivotal role in ensuring the reliability, performance, and scalability of our systems and services. You will collaborate closely with development teams to design and implement robust infrastructure solutions that support our growing product offerings. Your expertise will be essential in automating processes, monitoring system health, and troubleshooting complex issues, all while fostering a culture of reliability and efficiency within the organization.
Key facts
What you'll do
- Design, implement, and maintain scalable and reliable infrastructure solutions that meet the demands of our applications and services.
- Collaborate with software development teams to integrate reliability into the software development lifecycle, ensuring that new features are built with operational excellence in mind.
- Develop and manage monitoring and alerting systems to proactively identify and resolve issues before they impact users.
- Automate repetitive tasks and processes using scripting and configuration management tools to improve efficiency and reduce human error.
- Conduct post-mortem analyses of incidents to identify root causes and implement preventive measures, fostering a culture of continuous improvement.
- Participate in on-call rotations to provide support for production systems, ensuring high availability and quick resolution of incidents.
- Work closely with cross-functional teams to establish and enforce best practices for system reliability, performance, and security.
- Evaluate and implement new technologies and tools that enhance our infrastructure and improve our operational capabilities.
- Mentor junior engineers and contribute to their professional development, sharing your knowledge and expertise in site reliability engineering.
- Collaborate with product management to understand user requirements and translate them into technical solutions that enhance user experience.
- Prepare and present technical documentation and reports to stakeholders, ensuring transparency and alignment on reliability initiatives.
- Engage in capacity planning and performance tuning to ensure that our systems can handle growth and increased demand efficiently.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- A minimum of 5 years of experience in site reliability engineering, DevOps, or a related field, with a strong focus on infrastructure management.
- Proficiency in cloud platforms such as AWS, Google Cloud, or Azure, with hands-on experience in deploying and managing cloud-based applications.
- Strong scripting skills in languages such as Python, Bash, or Ruby, with experience in automation frameworks and tools.
- Familiarity with containerization technologies like Docker and orchestration tools such as Kubernetes.
- Experience with monitoring and logging tools like Prometheus, Grafana, ELK Stack, or similar technologies.
- Solid understanding of networking concepts, protocols, and security best practices.
- Excellent problem-solving skills and the ability to work effectively under pressure in a fast-paced environment.
- Strong communication skills, both written and verbal, with the ability to collaborate effectively with technical and non-technical stakeholders.
- A proactive mindset with a passion for improving system reliability and performance.
Nice to have
- Experience with infrastructure as code (IaC) tools such as Terraform or CloudFormation.
- Familiarity with CI/CD pipelines and tools like Jenkins, GitLab CI, or CircleCI.
- Knowledge of database management systems and performance tuning, particularly with SQL and NoSQL databases.
- Previous experience in a startup or high-growth environment, demonstrating adaptability and innovation.
- Contributions to open-source projects or involvement in the tech community.
Skills & tools
- Cloud platforms: AWS, Google Cloud, Azure
- Scripting languages: Python, Bash, Ruby
- Containerization: Docker, Kubernetes
- Monitoring tools: Prometheus, Grafana, ELK Stack
- Infrastructure as Code: Terraform, CloudFormation
- CI/CD tools: Jenkins, GitLab CI, CircleCI
Practical notes
- The specific location for this position is currently unknown, but candidates from various regions are encouraged to apply.
- This is a full-time role with competitive salary and benefits, including potential visa sponsorship for qualified candidates.
- Interested applicants can submit their applications through the provided link: [Bamboohr Careers](https://opench.bamboohr.com/careers/803).
Join Bamboohr and contribute to building reliable and scalable systems that enhance user experiences and drive innovation in our industry. Your expertise will be valued as we work together to create a more resilient infrastructure.