Cloud Operations Manager
panoptycRemoteFull Time2w ago
PythonRubyAWSKubernetesTerraformJenkinsCI/CDGitHub ActionsAIMLSecuritySOC
Job description
Cloud Operations Manager at panoptyc
About the role
We are looking for a Cloud Operations Manager to enhance our cloud, GPU compute, and edge device infrastructure. This role involves bringing structure to our AWS environment, establishing efficient CI/CD practices, and guiding our IT team to build dependable, scalable systems that support engineering progress.
Key facts
What you'll do
- Lead and develop IT operations personnel, offering technical direction, defining clear workflows, and improving their cloud engineering abilities through practical guidance.
- Manage and advance our AWS infrastructure, implementing sound architectural approaches, cost management strategies, security enhancements, and disaster recovery plans for web, AI/ML, and edge computing.
- Create and maintain automated CI/CD systems for infrastructure and applications, supporting deployments of containers, hardware-dependent services, and ML models.
- Enhance operational maturity by developing documentation, runbooks, change management protocols, incident response plans, and knowledge sharing procedures for cloud and edge systems.
- Collaborate with engineering teams to provide infrastructure that allows for rapid, secure deployments while ensuring system stability and security.
- Oversee incident resolution, conduct post-incident reviews, and implement preventative measures to minimize recurring issues.
Requirements
- Minimum 5 years of hands-on experience with AWS services such as EC2, RDS, S3, ECS, Fargate, IAM, infrastructure as code, networking, and security groups, including architecture and troubleshooting.
- Solid background in CI/CD tools like GitHub Actions, Jenkins, GitLab CI, or CircleCI, encompassing pipeline design, testing methodologies, and deployment automation.
- Proven experience in managing and developing technical team members, with the ability to patiently coach less experienced engineers through complex technical challenges.
- Demonstrated success in implementing and maintaining operational processes, including documentation standards, change management, on-call schedules, and incident response.
- Understanding of cloud security principles, IAM policies, SOC 2 considerations, and the use of infrastructure as code for audit purposes.
- Strong communication skills, capable of explaining technical details clearly to both engineering colleagues and non-technical individuals.
Nice to have
- Experience in optimizing costs for ML/GPU workloads.
- Proficiency with Terraform or CloudFormation for infrastructure as code.
- Familiarity with Kubernetes or other container orchestration platforms.
- Experience with monitoring and observability tools such as DataDog, CloudWatch, Grafana, or PagerDuty.
- Background in Site Reliability Engineering (SRE) practices.
- AWS certifications (e.g., Solutions Architect, SysOps Administrator).
- Experience with SOC 2 compliance and security audits.
- Scripting abilities in Python, Bash, or Ruby for automation tasks.
Skills & tools
- AWS
- CI/CD tools (GitHub Actions, Jenkins, GitLab CI, CircleCI)
- Infrastructure as Code (IaC)
- Containerization
- Incident Management
- Process Development
- Team Leadership
- Cloud Security
Practical notes
- Visa sponsorship is not provided for this role.
- Travel is not expected.
- Benefits information is not specified.
- Please submit your application to be considered.