
Staff Engineer Tech Lead Manager, ML Acceleration
Job description
About the role
We are seeking an experienced and visionary Principle Level Tech Lead Manager to build and lead our new Machine Learning (ML) Acceleration team. This role owns the end to end strategy, execution, and delivery of initiatives designed to accelerate ML model training across the organization. The primary mission is to drastically reduce the development cycle for new ML models and enable rapid hot-patching for issues within our deployed autonomous vehicle services. You will operate at the intersection of deep ML systems engineering and high impact people leadership. This position requires a hands on leader who blends technical depth in ML systems and performance optimization with strong management capabilities. You will recruit, mentor, and grow a high performing team, fostering a culture of innovation, collaboration, and continuous improvement.
Key facts
What you'll do
Define the technical vision and strategy for ML acceleration across the organization, setting clear goals and priorities aligned with company objectives.
Build, lead, and manage a high performing team of ML and infrastructure engineers focused on acceleration, providing technical guidance, mentorship, and career development.
Design, develop, and implement scalable and efficient ML acceleration solutions, including data pipeline optimization, large scale distributed training, data loader optimization, hardware acceleration, and model optimization techniques.
Identify and evaluate cutting edge technologies and methodologies to speed up ML training, conducting experiments to validate their impact.
Collaborate closely with ML research, ML training platform, and product teams to understand their needs and integrate acceleration solutions seamlessly into existing workflows.
Communicate complex technical concepts and strategies to both technical and non technical stakeholders, ensuring alignment and shared understanding.
Regularly measure and report on the impact of acceleration efforts, using data to drive decisions and continuously seek opportunities for further optimization and innovation.
Act as a technical expert and advocate for ML acceleration initiatives across the company, promoting best practices and raising the technical bar.
Own the end to end delivery of critical acceleration projects, from initial scoping and requirements gathering through design, implementation, testing, and rollout.
Foster a collaborative and inclusive team environment where engineers can thrive, share knowledge, and push the boundaries of ML systems performance.
Drive the adoption of MLOps principles and practices to streamline experimentation, deployment, and monitoring of accelerated ML workloads.
Evaluate hardware acceleration options, including GPUs, TPUs, and emerging accelerators, to maximize training throughput and efficiency.
Partner with infrastructure and platform teams to ensure that the underlying systems support high performance ML training at scale.
Continuously benchmark and profile ML workloads to identify bottlenecks and apply targeted optimizations that reduce training time and resource utilization.
Requirements
Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience.
8+ years of experience in software engineering, with at least 5+ years focused on Machine Learning systems and model performance optimization.
3+ years of experience in a technical lead or management role, with a proven track record of building and leading high performing teams.
Extensive experience with large scale ML model training and deployment, ideally in a production environment.
Strong understanding of distributed systems and cloud computing platforms such as AWS, GCP, and Azure.
Deep expertise in ML frameworks such as PyTorch or JAX.
Proficiency in performance profiling and optimization techniques for Deep Neural Networks.
Strong programming skills in Python; C++ experience is a plus.
Knowledge of MLOps principles and practices, including experiment tracking, model versioning, and deployment pipelines.
Excellent leadership, communication, and interpersonal skills.
Demonstrated ability to attract, hire, and retain top engineering talent in a competitive market.
Proven ability to drive complex technical projects from conception to completion, managing dependencies and delivering results on schedule.
Strong problem solving skills and a proactive, results oriented mindset.
Ability to thrive in a fast paced, dynamic environment and manage multiple priorities in a rapidly evolving organization.
Nice to have
C experience in performance critical components of ML frameworks or runtime systems.
Contributions to open source ML projects or engagement with the broader ML engineering community.
Experience with hardware specific optimizations, such as CUDA, oneAPI, or similar low level programming models.
Background in autonomous vehicle software stacks and safety critical systems.
Experience with containerization and orchestration platforms such as Kubernetes.
Practical notes
This is a full-time position based in the United States, with work locations in Boston, Massachusetts; Pittsburgh, Pennsylvania, and remote options within the U.S. Travel may be required between locations as needed. The role is eligible for standard company benefits as applicable. Candidates must be authorized to work in the United States without sponsorship for this position.