Vice President Site Reliability Engineering
Job description
About the role
In this pivotal position, you will lead the Site Reliability Engineering (SRE) team at Galaxy, focusing on the automation of infrastructure and ensuring the reliability of our systems. You will be responsible for shaping the vision of our automation roadmap while maintaining a hands-on approach to technical challenges. Your expertise will help us treat our infrastructure as a product, enhancing both engineering velocity and system stability. This role is crucial for managing complex hybrid environments and developing self-service platforms that support our engineering teams.
Key facts
What you'll do
- Lead a specialized SRE team dedicated to the design, implementation, and upkeep of automation tools and the systems they interact with.
- Establish and enforce standards for Infrastructure as Code (IaC) to guarantee consistent, repeatable, and secure deployments across the infrastructure ecosystem, with a strong emphasis on Terraform.
- Develop strategies for automated configuration and state management, optimizing Ansible playbooks and Packer image pipelines for Windows, Linux, and ESXi platforms.
- Oversee the monitoring and health of automation platforms, implementing Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to ensure high availability and performance of the tools that build servers.
- Manage the automated lifecycle of both physical and virtual assets, from initial template creation and deployment to automated patching, scaling, and decommissioning.
- Spearhead the creation of custom scripts and internal providers using languages like Python, Go, PowerShell, and Bash to enhance insights and tooling for our systems.
- Collaborate with other teams, particularly the Datacenter team, to facilitate workflows and address the collective needs of the organization.
- Analyze system behavior and resource utilization in virtual environments to optimize the performance of automated deployments.
- Provide mentorship and technical guidance to SREs, fostering a culture of automation-first and continuous improvement within the team.
Requirements
- A minimum of 6-10 years of experience in Infrastructure, SRE, or DevOps, with a strong focus on infrastructure automation at scale.
- Extensive knowledge of Terraform, including providers, modules, and state management, as well as Ansible for roles and playbooks.
- Proven hands-on experience with image creation tools such as Packer, Ansible, or SCCM for building standardized, hardened images in hybrid environments.
- Strong background in managing and automating virtual platforms like VMware (vSphere/vCenter) and cloud providers including Azure and AWS.
- Proficiency in scripting languages such as Python, Go, PowerShell, and Bash.
- Familiarity with observability tools like Splunk, ELK, Prometheus, or Grafana to monitor infrastructure health and automation metrics.
- Solid understanding of network topology and design, with experience in platforms such as Juniper Networks or Palo Alto.
- Expertise in Git, including branching strategies and pull request workflows, as well as CI/CD platforms like Jenkins, GitLab CI, or GitHub Actions.
- Comfortable managing, troubleshooting, and optimizing performance for both Windows Server and Linux environments.
Nice to have
- Previous experience in team leadership or management roles.
- Familiarity with Identity and Access Management (IAM) platforms such as Entra ID, Active Directory, or Okta.
- Experience with storage solutions, both block-based and object-based, whether on-premises (HP Alletra, EMC, DDN) or in the cloud (S3, Azure Blob).
- Knowledge of storage backup and disaster recovery administration using tools like Commvault or Veeam.
Skills & tools
- Terraform
- Ansible
- Packer
- VMware (vSphere/vCenter)
- Azure and AWS
- Python, Go, PowerShell, Bash
- Splunk, ELK, Prometheus, Grafana
- Git, Jenkins, GitLab CI, GitHub Actions
Practical notes
At Galaxy, we value diversity and are committed to providing equal employment opportunities to all employees and job applicants. We do not discriminate based on age, race, color, creed, religion, sex, gender identity, sexual orientation, marital status, national origin, disability, or any other characteristic protected by law. We strive to accommodate qualified applicants with disabilities unless it imposes an undue hardship on our operations. If you need assistance during the application process or require accommodations for an interview, please reach out to us at careers@galaxy.com.