AI Engineer, Model Distillation, Tesla AI
TeslaUSA3d ago
AIEngineeringremotecurated-jd
Job description
AI Engineer, Model Distillation, Tesla AI at Tesla.
About the role
Join the Tesla AI team to build intelligence for millions of robots by training frontier-scale foundation models. You will focus on transferring capabilities from massive teacher models to compact, efficient student models that operate under strict real-world constraints.
Key facts
What you'll do
- Create and run distillation pipelines to move reasoning, vision-language, and action policies from large teachers to student models for FSD, Optimus, and Digital Optimus.
- Develop distillation methods including logit-level, sequence-level, and trajectory-level techniques using synthetic data, soft-labels, SFT, and RL.
- Analyze scaling behaviors regarding teacher size, student capacity, compute allocation, and data mixtures to optimize on-robot performance.
- Engineer infrastructure for large-scale teacher inference, data curation, and distributed training while addressing memory and compute bottlenecks.
- Validate distilled models against teacher fidelity and production metrics to ensure performance and safety.
- Collaborate with cross-functional teams to deploy models into production environments.
- Build reusable tools and frameworks to support distillation across all Tesla AI model families.
Requirements
- Deep expertise in deep learning, specifically in compressing, training, or distilling large-scale multimodal, vision, or language models.
- Understanding of teacher-student dynamics and the balance between compute, latency, and capability.
- Knowledge of loss functions like KL or soft-targets and modern architectures such as hybrid attention or mixture of experts.
- Experience with post-training processes including RL, policy optimization, and supervised fine-tuning.
- Proficiency in distributed computing and large-scale training or inference systems.
- Strong Python programming skills and software engineering experience.
- Experience with JAX, TensorFlow, or PyTorch.
- Ability to troubleshoot system-level issues across training, data, and deployment.
Skills & tools
- Python
- PyTorch, TensorFlow, or JAX
- Distillation (Logit, sequence, trajectory)
- Distributed computing
- Supervised Fine-Tuning (SFT)
- Reinforcement Learning (RL)
Practical notes
Benefits begin on the first day of employment and include $0 payroll deduction medical, dental, and vision plans, HSA contributions, 401(k) with match, and stock purchase plans. Additional perks include family-building support, disability insurance, paid time off, and commuter benefits. Compensation is based on individual experience, skills, and market location.