Infrastructure and MLOps Engineer
Job description
About the role
Join the Software Infrastructure team to scale and manage systems that support our AI compute stack. You will build tools and services that streamline the development, testing, and deployment of Machine Learning software components. In this capacity, you will own the reliability and efficiency of the pipelines that empower our researchers and engineers. The role focuses on constructing the foundational platforms that allow complex AI workloads to run securely and at scale. You will act as a bridge between infrastructure capabilities and the daily needs of the data science and engineering teams. The work involves close collaboration with software developers to ensure that the deployment lifecycle is robust, observable, and maintainable. You will contribute directly to the productization of Machine Learning software by turning experimental workflows into production-grade services. This position is critical for ensuring that the infrastructure keeps pace with the innovation velocity of the organization.
Key facts
What you'll do
Develop and maintain services for AI research and engineering groups, ensuring high availability and performance.
Manage cloud infrastructure using Terraform to define, provision, and manage resources in a repeatable and scalable manner.
Deploy and operate services within Kubernetes and Docker environments, handling lifecycle and configuration complexities.
Support the productization of Machine Learning software by transforming prototypes into reliable, shared platforms.
Implement and refine CI/CD pipelines to automate testing, validation, and delivery of ML software components.
Monitor system health and performance, utilizing observability tools to identify and resolve incidents proactively.
Collaborate with cross-functional teams to gather requirements and translate them into robust infrastructure solutions.
Maintain and optimize the toolchain for ML orchestration, including frameworks such as NV Ray, KFP, or SkyPilot.
Manage interactions with ML accelerator hardware, utilizing interfaces and metrics from tools like DCGM for performance insights.
Contribute to the development of internal standards and best practices for infrastructure, security, and compliance.
Automate routine operational tasks to reduce manual overhead and improve the efficiency of the engineering workflows.
Engage in on-call duties to provide support for production incidents and ensure rapid resolution of critical issues.
Evaluate emerging technologies and assess their applicability to enhancing the existing AI compute stack.
Document architectural decisions, operational procedures, and service interfaces to ensure clarity and knowledge sharing.
Requirements
- Proficiency in Python is essential for developing automation scripts, building services, and interacting with APIs.
- Experience with Linux environments is required to manage servers, containers, and development workstations effectively.
- Familiarity with cloud providers such as AWS is necessary to design and operate resilient and cost-effective infrastructure.
- Understanding of CI/CD principles is mandatory to implement automated testing, builds, and deployments.
- Experience with Kubernetes is required to orchestrate containerized applications at scale.
- Practical knowledge in at least one of the following areas is required: maintaining ML applications, deploying ML orchestration tools (e.g., NV Ray, KFP, SkyPilot), or managing ML accelerator hardware (e.g., DCGM).
- Strong problem-solving skills are required to troubleshoot complex issues in distributed systems and networking.
- Excellent communication skills are necessary to work effectively with technical and non-technical stakeholders.
- The ability to work independently and as part of a collaborative team is essential in a fast-paced engineering environment.
- A commitment to following security and operational best practices is required to maintain a robust and resilient infrastructure.
Nice to have
Experience with Infrastructure as Code tools like Terraform or OpenTofu.
Proficiency with GitHub Actions for automating software workflows.
Knowledge of observability tools including Prometheus and Grafana.
Programming experience in Go, Java, or C++.
Practical notes
Benefits include a competitive salary, flexible working arrangements, generous annual leave, private medical insurance, a health cash plan, a dental plan, and a pension scheme with up to 5% matching. The package also includes life assurance, income protection, a parental leave policy, and an employee assistance program covering mental wellbeing and bereavement support.