Cloud Operations Manager
Job description
About the role
The is responsible for elevating our cloud, GPU compute, and edge device infrastructure to a more professional and scalable level. This position requires an individual who will take ownership of structuring our AWS environment and establishing robust, efficient CI/CD practices. The successful hire will act as a key guide for our IT team, helping them to build dependable systems that actively support engineering progress. The role focuses on enhancing operational maturity and ensuring that infrastructure is both stable and scalable. You will be expected to bring a high level of technical direction and practical guidance to the team. The position involves direct responsibility for the reliability and performance of our core infrastructure. Ultimately, this role is about enabling the engineering organization through strong operational leadership and process discipline.
Key facts
What you'll do
- Provide technical direction and coaching to IT operations personnel, defining workflows and elevating cloud engineering capabilities through hands-on mentorship.
- Architect, manage, and advance our AWS infrastructure with a focus on cost efficiency, security, reliability, and support for web, AI/ML, and edge computing workloads.
- Design, implement, and maintain automated CI/CD systems for both infrastructure and applications, handling containers, hardware-dependent services, and ML model deployments.
- Drive improvements in operational maturity by establishing comprehensive documentation, detailed runbooks, change management protocols, incident response plans, and knowledge transfer practices.
- Partner closely with engineering teams to deliver infrastructure solutions that enable rapid and secure deployments without compromising system stability or security.
- Lead incident resolution efforts, conduct thorough post-incident reviews, and implement preventative controls to reduce the likelihood of recurring issues.
- Define and enforce cloud security standards, including IAM policy design, SOC 2 considerations, and the application of infrastructure as code for auditability.
- Evaluate and integrate new tools and technologies such as container orchestration platforms and observability systems to enhance system reliability.
- Optimize infrastructure costs, with a particular focus on managing expenses related to ML and GPU-intensive workloads.
- Serve as a primary point of contact for infrastructure strategy, ensuring alignment with broader engineering and business objectives.
Requirements
- Possess a minimum of 5 years of hands-on experience with core AWS services including EC2, RDS, S3, ECS, Fargate, IAM, infrastructure as code, networking, and security groups, covering both architecture design and troubleshooting.
- Demonstrate a solid background in CI/CD tools such as GitHub Actions, Jenkins, GitLab CI, or CircleCI, including pipeline design, testing methodologies, and deployment automation.
- Show proven experience in managing and developing technical team members, with the ability to patiently coach less experienced engineers through complex technical problems.
- Have a demonstrated history of implementing and maintaining operational processes, including documentation standards, change management, on-call schedules, and incident response procedures.
- Exhibit a strong understanding of cloud security principles, IAM policies, SOC 2 compliance considerations, and the use of infrastructure as code for audit purposes.
- Communicate effectively with clarity, capable of explaining technical concepts to both engineering colleagues and non-technical stakeholders.
- Have direct experience with infrastructure as code tools and methodologies, ensuring that environments are versioned, repeatable, and secure.
- Maintain a strong attention to detail and a proactive approach to identifying potential infrastructure risks before they impact production.
Nice to have
- Experience in optimizing costs for ML/GPU workloads.
- Proficiency with Terraform or CloudFormation for infrastructure as code.
- Familiarity with Kubernetes or other container orchestration platforms.
- Experience with monitoring and observability tools such as DataDog, CloudWatch, Grafana, or PagerDuty.
- Background in Site Reliability Engineering (SRE) practices.
- AWS certifications (e.g., Solutions Architect, SysOps Administrator).
- Experience with SOC 2 compliance and security audits.
- Scripting abilities in Python, Bash, or Ruby for automation tasks.
Skills & tools
- AWS
- CI/CD tools (GitHub Actions, Jenkins, GitLab CI, CircleCI)
- Infrastructure as Code
- Containerization
- Incident Management
- Process Development
- Team Leadership
- Cloud Security
Practical notes
- Visa sponsorship is not provided for this role.
- Travel is not expected.
- Benefits information is not specified.
- Please submit your application to be considered.