
Cloud Operations Architect
Job description
About the role
Panoptyc is in search of a Cloud Operations Architect to strengthen our cloud infrastructure, particularly in GPU computing and edge devices. This position will be pivotal in leading our IT operations team, instilling operational rigor within our AWS ecosystem, and setting up best practices for continuous integration and continuous deployment (CI/CD) pipelines. You will play a crucial role in mentoring the IT team, ensuring the delivery of a reliable and scalable infrastructure that accelerates engineering efforts.
Key facts
What you'll do
- Provide guidance and mentorship to IT operations team members, offering technical insights and fostering their cloud engineering capabilities through practical coaching.
- Manage and enhance our AWS environment by employing appropriate architectural patterns, optimizing costs, reinforcing security measures, and ensuring robust disaster recovery protocols for web applications, AI/ML tasks, and edge computing solutions.
- Design and maintain automated CI/CD pipelines for both infrastructure and application services, ensuring support for containerized environments, hardware-specific deployments, and machine learning applications.
- Drive operational excellence by creating comprehensive documentation, runbooks, change management processes, incident response strategies, and facilitating knowledge transfer across cloud and edge systems.
- Collaborate closely with engineering teams to deliver infrastructure that promotes rapid and secure deployments while ensuring system reliability and security are upheld.
- Lead incident response initiatives, conduct blameless post-mortem analyses, and implement proactive measures to mitigate the recurrence of issues.
- Establish and enforce best practices for operational processes, ensuring they are effectively adopted within the team.
- Engage in continuous improvement efforts, identifying areas for enhancement within the cloud infrastructure and operational workflows.
- Evaluate new technologies and tools that could enhance our cloud operations and overall infrastructure resilience.
- Act as a liaison between technical and non-technical stakeholders, translating complex technical concepts into understandable terms.
Requirements
- A minimum of 5 years of experience working with AWS services, including EC2, RDS, S3, ECS, Fargate, IAM, infrastructure as code (IaC), networking, and security groups, with a strong focus on architecture and troubleshooting.
- Demonstrated expertise with CI/CD tools such as GitHub Actions, Jenkins, GitLab CI, or CircleCI, including experience in pipeline design, testing methodologies, and deployment automation.
- Proven track record in managing and developing technical team members, with a talent for coaching junior engineers through intricate technical challenges.
- Experience in implementing operational processes that are widely accepted, including documentation standards, change management protocols, on-call rotations, and incident response frameworks.
- Solid understanding of cloud security best practices, including IAM policies, SOC 2 compliance considerations, and infrastructure-as-code practices for maintaining audit trails.
- Excellent communication skills, with the ability to convey complex technical information effectively to both engineering teams and non-technical stakeholders.
Nice to have
- Experience in optimizing costs specifically for machine learning and GPU workloads.
- Familiarity with Terraform or CloudFormation for infrastructure as code.
- Knowledge of Kubernetes or other container orchestration technologies.
- Experience with monitoring and observability tools such as DataDog, CloudWatch, Grafana, or PagerDuty.
- Background in Site Reliability Engineering (SRE) methodologies.
- Relevant AWS certifications, such as Solutions Architect or SysOps Administrator.
- Experience with SOC 2 compliance and conducting security audits.
- Proficiency in scripting languages such as Python, Bash, or Ruby for automation purposes.
Skills & tools
- AWS (EC2, RDS, S3, ECS, Fargate, IAM, IaC, networking, security groups)
- CI/CD tools (GitHub Actions, Jenkins, GitLab CI, CircleCI)
- Terraform
- CloudFormation
- Kubernetes
- DataDog
- CloudWatch
- Grafana
- PagerDuty
- Python
- Bash
- Ruby
Practical notes
This position is fully remote, allowing you to work from anywhere in Latin America. As a full-time role, you will be expected to engage with the team regularly and contribute to our collaborative culture. While the salary details are not specified, we encourage candidates to inquire during the interview process. We are committed to fostering a diverse and inclusive workplace, and we welcome applicants from all backgrounds to apply.