Cloud Operations Architect
panoptycChileFull Time2d ago
PythonRubyAWSKubernetesTerraformJenkinsCI/CDGitHub ActionsAIMLSecuritySOC
Job description
Cloud Operations Architect at panoptyc.
About the role
panoptyc is looking for a Cloud Operations Architect to enhance our cloud, GPU compute, and edge device infrastructure. This role involves leading our IT operations team, bringing operational discipline to our AWS environment, and establishing best practices for CI/CD pipelines. You will mentor the IT team to deliver reliable, scalable infrastructure that supports engineering speed.
Key facts
What you'll do
- Guide and mentor IT operations team members, offering technical direction and developing their cloud engineering skills through hands-on coaching.
- Oversee and expand our AWS environment, applying proper architecture patterns, cost optimization, security hardening, and disaster recovery for web, AI/ML, and edge computing.
- Create and maintain automated CI/CD pipelines for infrastructure and application services, including support for containerized, hardware-dependent, and ML deployments.
- Build operational maturity through documentation, runbooks, change management, incident response, and knowledge transfer across cloud and edge systems.
- Collaborate with engineering teams to provide infrastructure that enables quick, secure deployments while maintaining system reliability and security.
- Lead incident response, conduct blameless post-mortems, and implement preventive measures to reduce recurring issues.
Requirements
- 5+ years working with AWS services such as EC2, RDS, S3, ECS, Fargate, IAM, IaC, networking, and security groups, with hands-on architecture and troubleshooting.
- Strong background with CI/CD tools like GitHub Actions, Jenkins, GitLab CI, or CircleCI, including pipeline design, testing strategies, and deployment automation.
- Experience managing and developing technical team members, with skill in coaching less experienced engineers through complex technical concepts.
- Proven ability to implement operational processes that are adopted, including documentation standards, change management, on-call rotations, and incident response.
- Understanding of cloud security best practices, IAM policies, SOC 2 considerations, and infrastructure-as-code for audit trails.
- Ability to explain complex technical information clearly to both engineering teams and non-technical stakeholders.
Nice to have
- Experience optimizing costs for ML/GPU workloads.
- Terraform or CloudFormation infrastructure-as-code experience.
- Knowledge of Kubernetes or other container orchestration.
- Experience with monitoring or observability tools like DataDog, CloudWatch, Grafana, or PagerDuty.
- Background in Site Reliability Engineering (SRE) practices.
- AWS certifications such as Solutions Architect or SysOps Administrator.
- Experience with SOC 2 compliance and security audits.
- Scripting skills in Python, Bash, or Ruby for automation.
Skills & tools
- AWS (EC2, RDS, S3, ECS, Fargate, IAM, IaC, networking, security groups)
- CI/CD tools (GitHub Actions, Jenkins, GitLab CI, CircleCI)
- Terraform
- CloudFormation
- Kubernetes
- DataDog
- CloudWatch
- Grafana
- PagerDuty
- Python
- Bash
- Ruby