SRE - Incidents & Monitoring
Job description
SRE - Incidents & Monitoring at Cialdnb.
About the role
Cialdnb is seeking a dedicated Site Reliability Engineer (SRE) specializing in incidents and monitoring to join our dynamic team. In this role, you will play a crucial part in ensuring the reliability and performance of our systems, focusing on incident management and proactive monitoring strategies. Your expertise will help us maintain high availability and enhance our operational efficiency, ultimately contributing to a experience for our users.
Key facts
What you'll do
- Collaborate with cross-functional teams to establish and refine incident response protocols, ensuring swift resolution of service disruptions.
- Develop and implement monitoring solutions that provide real-time insights into system performance and health, utilizing tools such as Prometheus, Grafana, or similar technologies.
- Analyze incident data to identify trends and root causes, facilitating continuous improvement in system reliability and performance.
- Design and maintain dashboards that visualize key performance indicators (KPIs) for various services, enabling teams to make informed decisions.
- Participate in on-call rotations to provide support during incidents, ensuring that critical systems are monitored and maintained around the clock.
- Create and update documentation related to incident management processes, monitoring configurations, and troubleshooting guides to enhance team knowledge.
- Conduct post-incident reviews to evaluate response effectiveness and implement recommendations for future improvements.
- Collaborate with development teams to integrate reliability best practices into the software development lifecycle, promoting a culture of shared responsibility for system reliability.
- Mentor junior engineers and provide guidance on incident management and monitoring best practices, fostering a learning environment within the team.
- Stay current with industry trends and emerging technologies related to site reliability engineering, continuously seeking opportunities to enhance our practices.
- Work closely with the security team to ensure that monitoring solutions comply with security policies and best practices.
- Assist in capacity planning and performance tuning efforts to ensure that systems can handle anticipated loads efficiently.
Requirements
- Proven experience as a Site Reliability Engineer or in a similar role, with a strong focus on incident management and monitoring.
- Proficiency in scripting languages such as Python, Bash, or Go, enabling automation of monitoring and incident response tasks.
- Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and their monitoring tools, demonstrating an understanding of cloud-native architectures.
- Experience with incident management tools (e.g., PagerDuty, Opsgenie) and monitoring solutions (e.g., Datadog, New Relic).
- Strong analytical skills, with the ability to interpret complex data sets and derive actionable insights.
- Excellent communication skills, capable of conveying technical information to both technical and non-technical stakeholders.
- A proactive mindset, with a commitment to continuous improvement and a passion for enhancing system reliability.
Nice to have
- Experience with container orchestration technologies such as Kubernetes or Docker, providing insights into modern deployment practices.
- Familiarity with configuration management tools (e.g., Ansible, Puppet, Chef) to streamline infrastructure management.
- Knowledge of DevOps practices and principles, contributing to a collaborative and efficient development environment.
- Understanding of networking concepts and protocols, aiding in troubleshooting and incident resolution.
- Previous experience in a fast-paced startup environment, demonstrating adaptability and a strong sense of ownership.
Skills & tools
- Monitoring and observability tools: Prometheus, Grafana, Datadog, New Relic
- Incident management platforms: PagerDuty, Opsgenie
- Scripting languages: Python, Bash, Go
- Cloud platforms: AWS, Azure, GCP
- Configuration management: Ansible, Puppet, Chef
- Containerization: Docker, Kubernetes
Practical notes
- This position is a contract role, and the specific duration will be discussed during the interview process.
- The exact salary for this role will be determined based on experience and qualifications.
- Cialdnb is open to candidates from various locations, and remote work is supported.
- Visa sponsorship may be available for qualified candidates, depending on the specific circumstances.
If you are about ensuring system reliability and thrive in a collaborative environment, we encourage you to apply for the SRE - Incidents & Monitoring position at Cialdnb. Join us in our mission to deliver exceptional service and reliability to our users.