AI DevOps / Cloud Engineer
Job description
AI DevOps / Cloud Engineer at Mex Digital.
About the role
You will design and own the entire cloud infrastructure foundation for the AI and Data function on AWS, starting from the ground up and evolving into a production-grade MLOps environment. This role requires you to build CI/CD pipelines that automate every release while ensuring they remain auditable, repeatable, and reliable. You will manage containerization and orchestration for all AI workloads, ensuring scalable and cost-efficient operations. You will serve as the general technology owner for the AI team, handling environment setup, access management, monitoring, and incident response in a fast-moving environment. This is a builder position that demands strong ownership and the ability to translate complex data science requirements into robust infrastructure. You will implement end-to-end observability and workflow orchestration to support model deployment and operational health. Additionally, you will manage integrations with key data and analytics platforms to centralize information within AWS. Finally, you will establish and maintain clear documentation and engineering standards that the growing AI team will follow.
Key facts
What you'll do
- Design, build, and own the full cloud infrastructure for the AI and Data function on AWS, including VPCs, IAM, networking, security groups, compute, storage, and cost management to ensure a solid foundation before any model goes near production.
- Build and maintain CI/CD pipelines for data engineers, ML engineers, and data scientists using GitHub Actions or GitLab CI, implementing GitOps principles and automated deployment workflows so that every release is automated, auditable, and repeatable.
- Own general technology operations for the AI team in the absence of a dedicated TechOps function, including environment setup, access management, developer tooling, system monitoring, incident response, and vendor and license management for all tools used by the team.
- Own Docker and Kubernetes across all AI workloads, building scalable, reliable container environments for model training, batch processing, and real-time inference while managing cluster health, resource allocation, and cost efficiency.
- Work closely with ML engineers to deploy models into production, building and maintaining model serving infrastructure, inference endpoints, and batch scoring pipelines while owning the deployment side of the ML lifecycle including packaging, versioning, rollout, and rollback strategies.
- Implement end-to-end observability across infrastructure, application performance, and ML model health, including alerting, dashboards, and on-call processes to ensure system reliability.
- Set up and manage workflow orchestration tools such as Airflow, Dagster, or Prefect, ensuring pipelines are reliable, retryable, and observable for complex data and ML operations.
- Evolve the MLOps practice over time by implementing drift detection, data quality checks, performance tracking, experiment tracking, and automated retraining triggers to improve model robustness and accuracy.
- Manage integrations from CDPs and product analytics platforms including Segment and Amplitude, as well as mobile attribution and engagement tools such as Adjust, Firebase, MoEngage, and JourneyFi into the central AWS data infrastructure.
- Support deployment, access management, and integration of metadata management and BI tools including OpenMetadata and Metabase within the cloud environment to enable data discovery and reporting.
- Ensure all cloud infrastructure and AI systems meet security and compliance standards, including secrets management, encryption, network security, and access controls to protect sensitive data and systems.
- Maintain clear, up-to-date documentation for all infrastructure, deployment processes, and operational runbooks while setting engineering standards that the AI team follows as it grows.
Requirements
- 5 to 10 or more years of experience in DevOps, Cloud Engineering, or Site Reliability Engineering with a proven track record of delivering reliable infrastructure.
- Proven experience building cloud infrastructure from scratch, not solely maintaining existing environments, demonstrating initiative and architectural ownership.
- Expert-level AWS skills are required and non-negotiable, covering a deep understanding of core services and best practices.
- Hands-on experience with Kubernetes in production environments, including cluster operations, networking, and troubleshooting.
- Experience with Infrastructure as Code using Terraform or CloudFormation to automate and version infrastructure deployments.
- Comfortable being the sole infrastructure owner on a team initially, with a strong ownership mindset and the ability to drive decisions independently.
- Able to work directly with data scientists and ML engineers to translate business and technical requirements into scalable infrastructure solutions.
- Experience in high-growth startups, scale-ups, or product companies is preferred, showing adaptability in fast-paced settings.
- Background in AI/ML infrastructure or MLOps is a strong advantage, aligning technical infrastructure with machine learning workflows.
- Azure familiarity is a plus, providing broader multi-cloud knowledge and flexibility.
Nice to have
- Experience with workflow orchestration tools such as Airflow, Dagster, or Prefect for building reliable and observable data pipelines.
- Background in AI/ML infrastructure or MLOps practices, including experiment tracking, model versioning, and deployment automation.
- Familiarity with metadata management and BI tools like OpenMetadata and Metabase for data discovery and governance.
- Experience with mobile attribution and analytics platforms such as Adjust, Firebase, MoEngage, and JourneyFi for integrating marketing and product data.
- Knowledge of security and compliance best practices for cloud infrastructure, including secrets management and encryption.
Practical notes
Full-time engagement based in the Dubai Office. Candidates must be available to work onsite during standard business hours. No visa sponsorship is provided by the company for this role. Candidates must ensure they have the right to work in the UAE or secure their own visa. The role requires immediate availability due to critical infrastructure needs. Travel is not required for this position.