B2B Senior Systems Site Reliability Engineer
Job description
About the role
In this B2B Senior Systems Site Reliability Engineer role at Jamf, you will play a crucial role in ensuring the reliability and performance of our B2B systems. You will collaborate with cross-functional teams to enhance system availability and efficiency while addressing operational challenges. You will own the design and execution of strategies that guarantee high availability for critical B2B services. This position requires you to drive initiatives that streamline incident response and post-incident review processes. You will be responsible for bridging the gap between development velocity and infrastructure stability. Your work will directly influence the reliability roadmap for enterprise-facing platforms. You will contribute to the creation of runbooks that standardize operational procedures across the team. This role demands a proactive mindset to identify potential failures before they impact business operations. You will architect and maintain resilient infrastructure for B2B systems, ensuring scalability and robustness under varying loads. You will implement automated monitoring solutions to detect anomalies and trigger alerts before service degradation occurs. You will lead investigations into complex system outages, utilizing root cause analysis to prevent future recurrences. You will optimize deployment pipelines to enhance the efficiency of CI/CD workflows and reduce lead time for changes. You will coordinate with development teams to integrate reliability principles throughout the software development lifecycle. You will evaluate and recommend new technologies and tools that improve the efficiency of cloud resource utilization. You will establish performance benchmarks and conduct regular reviews to ensure systems meet established service level objectives. You will manage on-call schedules and provide technical leadership during critical incidents affecting customer-facing services. You will develop and maintain detailed documentation for system architectures and operational procedures. You will collaborate with security teams to ensure that infrastructure changes comply with organizational security policies. You will streamline log aggregation and analysis processes to improve visibility into system behavior. You will mentor junior engineers on best practices for system observability and troubleshooting techniques. You will drive the adoption of infrastructure-as-code methodologies to standardize environment provisioning. You will participate in capacity planning exercises to forecast future resource needs and prevent bottlenecks.
Key facts
What you'll do
- Architect and maintain resilient infrastructure for B2B systems, ensuring scalability and robustness under varying loads.
- Implement automated monitoring solutions to detect anomalies and trigger alerts before service degradation occurs.
- Lead investigations into complex system outages, utilizing root cause analysis to prevent future recurrences.
- Optimize deployment pipelines to enhance the efficiency of CI/CD workflows and reduce lead time for changes.
- Coordinate with development teams to integrate reliability principles throughout the software development lifecycle.
- Evaluate and recommend new technologies and tools that improve the efficiency of cloud resource utilization.
- Establish performance benchmarks and conduct regular reviews to ensure systems meet established service level objectives.
- Manage on-call schedules and provide technical leadership during critical incidents affecting customer-facing services.
- Develop and maintain detailed documentation for system architectures and operational procedures.
- Collaborate with security teams to ensure that infrastructure changes comply with organizational security policies.
- Streamline log aggregation and analysis processes to improve visibility into system behavior.
- Mentor junior engineers on best practices for system observability and troubleshooting techniques.
- Drive the adoption of infrastructure-as-code methodologies to standardize environment provisioning.
- Participate in capacity planning exercises to forecast future resource needs and prevent bottlenecks.
Requirements
- Minimum of 5 years of experience in a Site Reliability Engineering or similar role managing production infrastructure.
- Proficiency in cloud platforms such as AWS, Azure, or Google Cloud, with demonstrated ability to design solutions in these environments.
- Strong knowledge of container orchestration tools like Kubernetes, including cluster operations and lifecycle management.
- Experience with CI/CD pipelines and automation tools to enable reliable and repeatable deployment processes.
- Familiarity with monitoring and logging tools such as Prometheus or ELK stack for real-time system insights.
- Solid understanding of networking concepts and protocols, including TCP/IP, DNS, and load balancing mechanisms.
- Bachelor's degree in Computer Science, Engineering, or a related field that provides a foundational understanding of computing principles.
Nice to have
- Experience with scripting languages like Python or Bash to automate operational tasks and improve workflow efficiency.
- Knowledge of security best practices in cloud environments, including identity management and network segmentation.
- Familiarity with configuration management tools like Ansible or Terraform to automate infrastructure provisioning.
Practical notes
This position is open to candidates requiring visa sponsorship. The engagement is full-time, and remote work is available from Poland. Benefits include flexible working arrangements and a comprehensive health package. Applications will be accepted until the position is filled, and candidates are encouraged to submit detailed information regarding their operational experience. There are no specific travel requirements associated with this role, and the hiring process will include technical interviews focusing on real-world scenario problem-solving.