DevOps Engineer
Job description
About the role
You will own the design, implementation, and operation of Later's cloud infrastructure and DevOps toolchain with a focus on reliability, automation, and security. You will partner closely with product, data, and machine learning teams to deliver scalable platforms that unblock fast and safe experimentation. You will play a key role in extending our Kubernetes and GitOps foundations while contributing to emerging MLOps capabilities that power model deployment and monitoring. You will help define and drive infrastructure standards that balance velocity with operational robustness across engineering and data organizations. You will act as a key collaborator in cross-functional initiatives that bridge infrastructure, data science, and product delivery.
Key facts
What you'll do
Support the development and execution of the infrastructure roadmap across DevOps and MLOps in partnership with product and data growth plans.
Partner with engineering, data, and ML teams to ensure scalability, reliability, security, and automation are embedded in application infrastructure and machine learning workflows.
Evaluate and adopt DevOps and MLOps tools that improve system efficiency, observability, developer experience, model deployment, and operational reliability.
Contribute to platform standards that make infrastructure, CI/CD, data pipelines, and ML systems more repeatable, secure, and easier to operate for the entire organization.
Support the evolution of cloud-native practices that enable faster product delivery while preparing the platform for future AI and ML initiatives.
Build and manage infrastructure to deploy ML models into production reliably using CI/CD pipelines, Flask-based APIs, and orchestration tools such as Airflow, Kubeflow, or Argo Workflows.
Automate training pipelines, model registry, validation, deployment, and rollback strategies using tools such as Amazon SageMaker Interface and Postman for testing.
Build systems to monitor model performance, latency, data drift, and resource usage using Amazon CloudWatch, Prometheus, and Grafana.
Design and maintain tools and systems to support model versioning, experiment tracking using tools like MLflow and Amazon SageMaker Studio Notebooks, and reproducible training workflows.
Operate across GCP and AWS to manage training and inference infrastructure, BigQuery datasets, and GPU workloads in production environments.
Use infrastructure as code tools such as Terraform or CloudFormation to manage cloud infrastructure in a scalable and repeatable manner.
Work with Data Scientists, Analysts, Platform Engineers, and Product Engineers to support their end-to-end ML workflows and remove operational bottlenecks.
Implement robust deployment patterns, observability practices, and incident response processes to ensure high availability and quick recovery.
Collaborate on cost optimization strategies for cloud resources while maintaining performance and reliability for critical workloads.
Requirements
Demonstrate a strong foundation in DevOps practices and cloud infrastructure operations with a proven track record of delivering reliable systems.
Bring experience with AWS services, Kubernetes, and container orchestration, along with a solid understanding of networking, security, and compliance fundamentals.
Show proficiency in infrastructure as code using Terraform or CloudFormation and experience with CI/CD pipelines and GitOps workflows.
Provide evidence of working with containerized applications, Docker, and orchestration platforms such as Kubernetes in production or near-production environments.
Have hands-on experience with monitoring, logging, and observability tools such as CloudWatch, Prometheus, or Grafana in live systems.
Demonstrate scripting and automation skills in at least one high-level programming or scripting language for operational tasks and pipeline development.
Bring experience with MLOps concepts, model deployment strategies, and tooling such as SageMaker, MLflow, or similar platforms for managing ML lifecycles.
Show collaboration experience working with data teams, data scientists, and product engineers to deliver shared infrastructure and workflows.
Nice to have
Preferred experience with GCP services, BigQuery, and GPU-intensive workloads in cloud environments.
Familiarity with Postman for API testing of model endpoints and SageMaker interfaces.
Experience contributing to platform teams and establishing standards that scale across multiple engineering and data teams.
Practical notes
This role operates across GCP and AWS to manage training/inference infrastructure, BigQuery datasets, and GPU workloads.