Senior Site Reliability Engineer
Job description
About the role
As a , you will play a pivotal role in ensuring the reliability and performance of our systems and services. Your expertise will help us maintain high availability while optimizing our infrastructure for scalability and efficiency. This position is ideal for someone who thrives in a dynamic environment and is passionate about implementing best practices in site reliability engineering. You will collaborate closely with development teams to enhance the overall user experience and contribute to the continuous improvement of our systems.
Key facts
What you'll do
- Design, implement, and maintain scalable and reliable infrastructure solutions that support Vistairhr's applications and services.
- Monitor system performance and reliability, proactively identifying and resolving issues before they impact users.
- Collaborate with software development teams to integrate reliability into the software development lifecycle, ensuring that new features are designed with operational excellence in mind.
- Develop automation tools and scripts to streamline operational processes, reducing manual intervention and increasing efficiency.
- Conduct post-incident reviews to identify root causes and implement preventive measures to avoid future occurrences.
- Manage and optimize cloud infrastructure, ensuring cost-effectiveness and performance efficiency.
- Implement and maintain CI/CD pipelines to facilitate rapid and reliable software deployments.
- Participate in on-call rotations, providing support for production systems and responding to incidents as they arise.
- Create and maintain comprehensive documentation for systems, processes, and procedures to ensure knowledge sharing across the team.
- Stay current with industry trends and emerging technologies, evaluating their potential application within Vistairhr.
- Mentor junior engineers and contribute to their professional development, fostering a culture of learning and collaboration.
- Engage in capacity planning and performance tuning to ensure systems can handle future growth and demand.
Requirements
- A minimum of 5 years of experience in site reliability engineering, DevOps, or a related field.
- Proficiency in cloud platforms such as AWS, Azure, or Google Cloud, with hands-on experience in deploying and managing cloud-based applications.
- Strong knowledge of containerization technologies like Docker and orchestration tools such as Kubernetes.
- Experience with monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack) to ensure system health and performance.
- Familiarity with scripting languages such as Python, Bash, or Go for automation and tool development.
- Understanding of networking concepts and protocols, including TCP/IP, DNS, and HTTP/S.
- Excellent problem-solving skills and the ability to work under pressure in a fast-paced environment.
- Strong communication skills, both written and verbal, with the ability to collaborate effectively with cross-functional teams.
- A proactive mindset with a focus on continuous improvement and operational excellence.
Nice to have
- Experience with configuration management tools such as Ansible, Puppet, or Chef.
- Knowledge of database technologies (SQL and NoSQL) and their performance optimization.
- Familiarity with security best practices and compliance standards in cloud environments.
- Previous experience in a startup or fast-growing company, adapting to changing requirements and priorities.
- Contributions to open-source projects or a strong presence in the tech community.
Skills & tools
- Cloud platforms: AWS, Azure, Google Cloud
- Containerization: Docker, Kubernetes
- Monitoring: Prometheus, Grafana, ELK stack
- Scripting: Python, Bash, Go
- Configuration management: Ansible, Puppet, Chef
- CI/CD tools: Jenkins, GitLab CI, CircleCI
Practical notes
- This position is based in Barcelona, Catalonia, and may require occasional travel for team meetings or conferences.
- Candidates should be prepared for a thorough interview process that includes technical assessments and behavioral interviews.
- Vistairhr is committed to fostering a diverse and inclusive workplace, and we encourage applications from individuals of all backgrounds.
- For more information about the company culture and values, please visit our website or reach out to current employees on LinkedIn.
META
Company: Vistairhr
Title: Senior Site Reliability Engineer
Listed
location: Barcelona, Catalonia
Job type: Full-time
Apply URL: https://vistairhr.bamboohr.com/careers/219
Tags: Engineering, Reliability