Site Reliability Engineer Lead
DropsuiteBandung CityFull Time
Job description
About the role
As a , you will play a role in ensuring the reliability, performance, and scalability of our cloud-based services. This position is ideal for an experienced engineer who is about building and maintaining systems while leading a team of talented engineers. You will collaborate closely with cross-functional teams to enhance our infrastructure and improve our service delivery, ensuring that our customers receive the highest quality of service.
Key facts
What you'll do
- Lead a team of site reliability engineers, providing mentorship and guidance to foster professional growth and technical excellence.
- Design and implement scalable and reliable infrastructure solutions that support Dropsuite's cloud services.
- Develop and maintain monitoring and alerting systems to ensure the health and performance of our applications and infrastructure.
- Collaborate with software engineering teams to enhance the reliability and performance of applications through best practices in deployment and operations.
- Automate operational processes to improve efficiency and reduce manual intervention, using tools such as Terraform, Ansible, or similar.
- Conduct post-mortem analyses of incidents to identify root causes and implement preventive measures to avoid recurrence.
- Establish and enforce service level objectives (SLOs) and service level agreements (SLAs) to maintain high service quality.
- Participate in on-call rotations to provide support for production systems and respond to incidents as they arise.
- Evaluate and recommend new technologies and tools that can enhance our infrastructure and operational capabilities.
- Collaborate with product teams to understand customer needs and translate them into reliable and scalable solutions.
- Drive a culture of continuous improvement by implementing best practices in site reliability engineering and DevOps methodologies.
- Prepare and present reports on system performance, reliability metrics, and team progress to stakeholders and management.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field.
- Minimum of 5 years of experience in site reliability engineering, DevOps, or a related discipline.
- Strong expertise in cloud platforms such as AWS, Google Cloud, or Azure.
- Proficient in scripting and programming languages such as Python, Go, or Bash.
- Experience with containerization technologies like Docker and orchestration tools such as Kubernetes.
- Solid understanding of networking, security, and database management principles.
- Familiarity with CI/CD pipelines and tools like Jenkins, GitLab CI, or CircleCI.
- Excellent problem-solving skills and the ability to work under pressure in a fast-paced environment.
- Strong communication skills, with the ability to convey complex technical concepts to non-technical stakeholders.
- Proven experience in leading and mentoring teams, fostering a collaborative and innovative work environment.
Nice to have
- Experience with observability tools such as Prometheus, Grafana, or ELK stack.
- Knowledge of configuration management tools like Puppet or Chef.
- Familiarity with agile methodologies and experience working in an agile environment.
- Previous experience in a startup or fast-growing company is a plus.
- Understanding of compliance and regulatory standards relevant to cloud services.
Skills & tools
- Cloud Platforms: AWS, Google Cloud, Azure
- Programming Languages: Python, Go, Bash
- Containerization: Docker, Kubernetes
- CI/CD Tools: Jenkins, GitLab CI, CircleCI
- Monitoring Tools: Prometheus, Grafana, ELK stack
- Configuration Management: Terraform, Ansible, Puppet, Chef
Practical notes
- This position is based in Bandung City, West Java, and may require occasional travel for team meetings or conferences.
- The role offers opportunities for professional development and growth within the company, including access to training and certification programs.
- Candidates should be prepared to participate in technical interviews and provide examples of past work or projects relevant to site reliability engineering.
- Dropsuite is committed to fostering a diverse and inclusive workplace, and we encourage applications from all qualified individuals.
META
Company: Dropsuite
Title: Site Reliability Engineer Lead
Listed
location: Bandung City, West Java
Job type: Full-time
Apply URL: https://dropsuite.bamboohr.com/careers/228
Tags: Engineering, Reliability