Senior Site Reliability Engineer
Job description
About the role
Accelbyte is seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team in Sleman, Yogyakarta. In this pivotal role, you will be responsible for ensuring the reliability, availability, and performance of our gaming infrastructure. As a Senior SRE, you will collaborate closely with development teams to implement best practices in site reliability and foster a culture of continuous improvement. Your expertise will be crucial in maintaining the seamless operation of our services, allowing gamers around the world to enjoy a flawless gaming experience.
Key facts
What you'll do
- Design and implement robust monitoring and alerting systems to proactively identify and resolve issues before they impact users.
- Collaborate with software engineering teams to develop and maintain scalable and reliable applications, ensuring high availability and performance.
- Conduct post-mortem analyses of incidents to identify root causes and implement preventive measures, fostering a culture of learning and improvement.
- Automate operational processes using scripting and configuration management tools to enhance efficiency and reduce manual intervention.
- Manage and optimize cloud infrastructure, ensuring cost-effectiveness and resource utilization while maintaining performance standards.
- Develop and maintain documentation related to system architecture, operational procedures, and incident response protocols.
- Participate in on-call rotations to provide support for production systems, ensuring rapid response to incidents and minimizing downtime.
- Evaluate and implement new technologies and tools that can enhance system reliability and operational efficiency.
- Mentor junior engineers and share knowledge on best practices in site reliability engineering, fostering a collaborative learning environment.
- Work closely with cross-functional teams to align on service level objectives (SLOs) and service level agreements (SLAs) that meet business needs.
- Conduct capacity planning and performance tuning to ensure systems can handle anticipated loads and scale effectively.
- Engage in continuous improvement initiatives to enhance the overall reliability and efficiency of our systems and processes.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- A minimum of 5 years of experience in site reliability engineering, DevOps, or a related discipline.
- Proficiency in cloud platforms such as AWS, Azure, or Google Cloud, with hands-on experience in managing cloud infrastructure.
- Strong knowledge of containerization technologies like Docker and orchestration tools such as Kubernetes.
- Experience with monitoring and logging tools such as Prometheus, Grafana, ELK Stack, or similar technologies.
- Solid understanding of networking concepts, protocols, and security best practices.
- Proficient in scripting languages such as Python, Bash, or Go, with experience in automation and configuration management tools like Terraform or Ansible.
- Excellent problem-solving skills and the ability to work under pressure in a fast-paced environment.
- Strong communication skills, with the ability to convey technical concepts to non-technical stakeholders.
Nice to have
- Experience in the gaming industry or familiarity with game development processes.
- Knowledge of microservices architecture and API design principles.
- Familiarity with CI/CD pipelines and tools such as Jenkins, GitLab CI, or CircleCI.
- Experience with database management systems, both SQL and NoSQL.
- Understanding of incident management frameworks like ITIL or similar methodologies.
Skills & tools
- Cloud Platforms: AWS, Azure, Google Cloud
- Containerization: Docker, Kubernetes
- Monitoring: Prometheus, Grafana, ELK Stack
- Scripting: Python, Bash, Go
- Automation: Terraform, Ansible
- CI/CD: Jenkins, GitLab CI, CircleCI
Practical notes
This position is based in Sleman, Yogyakarta, and is a permanent role with competitive salary and benefits. Accelbyte is committed to fostering a diverse and inclusive workplace. We support relocation and visa sponsorship for qualified candidates. If you are passionate about site reliability and eager to contribute to the gaming industry, we encourage you to apply through our careers page at [Accelbyte Careers](https://accelbyte.bamboohr.com/careers/25). We look forward to welcoming a new member to our team who shares our dedication to delivering exceptional gaming experiences.