Software Engineer - Site Reliability Engineer
Job description
About the role
As a Software Engineer specializing in Site Reliability Engineering (SRE) at Alkira, you will play a pivotal role in ensuring the reliability, availability, and performance of our cloud networking solutions. This position requires a blend of software engineering and systems engineering skills to build and maintain scalable systems that can handle the demands of our rapidly growing customer base. You will collaborate with cross-functional teams to enhance our infrastructure, automate processes, and improve the overall user experience. Your contributions will directly impact the efficiency and effectiveness of our services, making this role both challenging and rewarding.
Key facts
What you'll do
- Design, implement, and maintain robust, scalable systems that support Alkira's cloud networking services.
- Collaborate with software development teams to integrate reliability into the software development lifecycle, ensuring that new features are reliable and performant.
- Monitor system performance and reliability metrics, proactively identifying and resolving issues before they impact customers.
- Develop and implement automation tools and scripts to streamline operational processes and reduce manual intervention.
- Create and maintain documentation related to system architecture, processes, and operational procedures to ensure knowledge sharing across teams.
- Participate in on-call rotations, providing support for production incidents and ensuring timely resolution of issues.
- Conduct post-mortem analyses for incidents, identifying root causes and implementing preventive measures to avoid future occurrences.
- Work closely with product and engineering teams to design and implement disaster recovery and business continuity plans.
- Evaluate and recommend new technologies and tools that can enhance system reliability and performance.
- Engage in capacity planning and performance tuning to ensure that systems can handle growth and increased demand.
- Collaborate with security teams to ensure that systems are secure and compliant with industry standards.
- Foster a culture of reliability and continuous improvement within the engineering organization.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- A minimum of 3 years of experience in software engineering or site reliability engineering roles.
- Proficiency in programming languages such as Python, Go, or Java, with a strong understanding of software development principles.
- Experience with cloud platforms (e.g., AWS, Azure, GCP) and container orchestration technologies (e.g., Kubernetes, Docker).
- Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack) to track system performance and reliability.
- Strong problem-solving skills and the ability to troubleshoot complex systems in a fast-paced environment.
- Experience with CI/CD pipelines and version control systems (e.g., Git) to facilitate efficient software delivery.
- Knowledge of networking concepts and protocols, including TCP/IP, DNS, and HTTP/S.
- Excellent communication skills, with the ability to work collaboratively across teams and articulate technical concepts to non-technical stakeholders.
Nice to have
- Experience with Infrastructure as Code (IaC) tools such as Terraform or CloudFormation.
- Familiarity with configuration management tools like Ansible, Puppet, or Chef.
- Understanding of security best practices in cloud environments and experience with security tools.
- Previous experience in a startup environment, demonstrating adaptability and a proactive approach to challenges.
- Contributions to open-source projects or participation in tech communities related to SRE or cloud computing.
Skills & tools
- Programming Languages: Python, Go, Java
- Cloud Platforms: AWS, Azure, GCP
- Containerization: Docker, Kubernetes
- Monitoring Tools: Prometheus, Grafana, ELK stack
- CI/CD: Jenkins, GitLab CI, CircleCI
- Configuration Management: Ansible, Puppet, Chef
- Infrastructure as Code: Terraform, CloudFormation
Practical notes
- The specific location for this role is currently unknown.
- Salary information is not provided; however, it is expected to be competitive based on industry standards.
- Visa sponsorship may be available for qualified candidates, but this will need to be confirmed during the application process.
- Interested candidates can apply through the Alkira careers page at https://alkira.bamboohr.com/careers/224.
Join Alkira and be part of a dynamic team that is redefining cloud networking solutions. Your expertise in site reliability will help us deliver exceptional services to our customers while fostering a culture of innovation and excellence.