Senior Software Engineer, Machine Learning Infrastructure
Job description
Senior Software Engineer, Machine Learning Infrastructure at Match Group.
About the role
This position is responsible for the end-to-end lifecycle of machine learning infrastructure at a global scale. The successful candidate will design, construct, and maintain the foundational platforms that enable machine learning experimentation, training, deployment, and monitoring. The focus is on building robust systems that process large-scale datasets, including hundreds of billions of data points, across multiple business units. The role requires a balance of platform engineering, system optimization, and cross-functional collaboration to ensure technical excellence and operational reliability. You will partner closely with data scientists and engineers to translate complex ML requirements into scalable infrastructure solutions.
Key facts
What you'll do
Architect and sustain the infrastructure that powers machine learning workflows, including data processing and moderation pipelines designed to handle massive data volumes while integrating seamlessly with trust and safety operations.
Oversee the deployment and management of production ML systems using internal deployment tools to optimize compute and storage resources for reliability, scalability, and cost efficiency.
Design and develop application programming interfaces, including REST, gRPC, and GraphQL, to support internal ML platform services and system integrations.
Utilize observability tools to monitor deployment and system performance, ensuring strict adherence to technical specifications and service-level objectives.
Develop and implement model evaluation, validation, and quality assurance processes, such as A/B testing frameworks and automated evaluation systems, to guarantee model accuracy and performance.
Design and maintain scalable ML platform systems and data infrastructure using distributed data technologies, including Apache Spark, Kafka, Flink, and Databricks to support global data processing and analytics needs.
Analyze infrastructure requirements and design technical solutions within defined scalability, performance, and cost constraints for data processing and analytics.
Support the full ML lifecycle infrastructure, including model training, serving, monitoring, feature stores, and evaluation systems, with a strong emphasis on platform engineering and self-service capabilities.
Mentor junior engineers on ML systems, backend systems, scalable data pipelines, production reliability, and deployment best practices to elevate team capabilities.
Participate in hiring activities by conducting technical interviews and providing input on candidate evaluations to build high-performing engineering teams.
Requirements
Candidates must possess a Bachelor's degree or its U.S. equivalent in Computer Science, Computer Engineering, or a related field, plus 5 years of professional experience as a Machine Learning Engineer, Site Reliability Engineer, or in a role performing ML infrastructure or backend software engineering.
In lieu of a Bachelor's degree, a Master's degree or U.S. equivalent in a related field is acceptable with 3 years of relevant professional experience.
You must demonstrate 3 years of professional experience designing and implementing large-scale distributed ML platform systems using big data technologies such as Apache Spark, Apache Kafka, Apache Flink, or Databricks.
This requires 3 years of experience using multiple modern programming languages, including Python, Scala, Java, or Go, to develop ML platform systems, backend services, data processing jobs, and automation tools supporting the ML lifecycle.
Required experience includes 2 years of professional work with modern cloud platforms, including AWS, Azure, or GCP, and utilizing infrastructure-as-code practices, containerization tools such as Docker on managed orchestration platforms including Amazon EKS or Amazon ECS, and monitoring systems based on Prometheus metrics and Grafana dashboards.
Candidates must have 2 years of experience designing and building infrastructure for recommendation systems, moderation pipelines, or large language model serving and deployment systems, including experience with modern ML serving frameworks such as Ray Serve or Triton.
Additional requirements include 2 years of professional experience in large-scale database design and optimization, and data pipeline performance tuning to support efficient data access patterns for ML workflows, including working with analytical storage systems including Delta Lake or data warehouses.
You must have 1 year of professional experience leading technical initiatives across multiple engineering teams, including establishing platform ownership models and driving adoption of shared ML infrastructure components.
1 year of professional experience is required in designing and implementing CI/CD automation pipelines and GitOps practices for ML infrastructure, using tools including Terraform, Terragrunt, Helm, and internal GitOps systems, together with continuous integration systems to manage deployment strategies.