ML Infrastructure Engineer
Job description
About the role
This position invites you to join the software engineering infrastructure team responsible for building and maintaining high-performance, globally distributed systems that power critical workloads. You will own the end-to-end model delivery pipeline, encompassing training workflows, serving architectures, and optimization strategies for our advertising ecosystem. The role requires deep collaboration with backend and research science teams to translate ambitious product goals into scalable technical roadmaps and implementations. You will act as a key technical influence, shaping infrastructure decisions that impact reliability, latency, and efficiency across the organization. As a mentor on the team, you will elevate the technical capabilities of other engineers through guidance, code reviews, and architectural leadership. This position is centered on designing systems that handle massive scale while ensuring robustness and performance under demanding conditions. You will have ownership over the full lifecycle of infrastructure components, from initial design through deployment and long-term operations. The work involves solving challenging distributed systems problems where infrastructure quality directly affects business outcomes and user experience.
Key facts
What you'll do
- Architect and implement large-scale distributed systems that form the backbone of global infrastructure.
- Partner with backend and research science teams to define and refine product and technology roadmaps.
- Drive initiatives to improve online model performance and streamline the model delivery pipeline.
- Engage with multiple engineering groups to resolve complex technical challenges that span domains.
- Provide mentorship and technical direction to team members to enhance overall engineering excellence.
- Design and optimize systems for high throughput and low latency in production environments.
- Evaluate and integrate new technologies to strengthen the capabilities of the model delivery pipeline.
- Lead projects that require cross-functional coordination and strategic technical decision-making.
- Ensure infrastructure components meet stringent reliability, scalability, and security standards.
- Contribute to the development of tools that enable other teams to operate efficiently at scale.
- Analyze performance bottlenecks and implement solutions that enhance system efficiency and user experience.
- Collaborate on the design of serving architectures that support dynamic workloads and diverse model types.
- Participate in on-call rotations to address critical incidents and maintain high availability.
- Document architectural decisions and operational procedures to support long-term maintainability.
- Work on proof-of-concept implementations to validate technical approaches before full-scale rollout.
Requirements
- Hold 0-2 years of professional experience in relevant roles within software infrastructure or similar domains.
- Possess a BS and/or MS degree in Computer Science from an accredited institution.
- Demonstrate a strong grasp of computer science fundamentals, including algorithms, data structures, and coding practices.
- Show proven ability to independently initiate and manage projects from conception to completion.
- Exhibit strong problem-solving skills and the capacity to debug complex issues in distributed systems.
- Communicate effectively with both technical and non-technical stakeholders to align on objectives and priorities.
- Maintain a high level of ownership and accountability for assigned tasks and project outcomes.
- Adhere to best practices in software engineering, including version control, testing, and code quality.
Nice to have
- Demonstrate proficiency in C++, Python, and/or Golang through prior projects or professional experience.
Skills & tools
- Distributed systems architecture forms the foundation of your technical approach and design decisions.
- Experience with machine learning infrastructure, including training and serving pipelines, is highly relevant.
- Strong coding skills in C++ enable you to build performance-critical components efficiently.
- Proficiency in Python supports rapid development and integration with data science workflows.
- Familiarity with Golang helps in building concurrent and scalable backend services.
Practical notes
- Compensation includes equity eligibility, medical, dental, vision, life, and disability insurance.
- Benefits include a 401(k) plan, unlimited discretionary time off, 10 paid holidays, and 80 hours of paid sick leave annually.
- The application window is expected to close within 30 days of the posting date.
- Apply online via the company portal.
- Direct inquiries to peopleops@applovin.com.