Senior Research Engineer, LLM Training & Post-Training
Job description
About the role
You will own the design and execution of large language model training and post-training pipelines within Lightning AI's production research infrastructure. You will translate research insights into scalable training systems and optimize model quality, throughput, and reliability across distributed workloads. You will collaborate closely with researchers, infrastructure engineers, and customers to align model behavior with real-world deployment constraints. You will implement and evaluate advanced post-training techniques such as supervised fine-tuning, reinforcement learning, and preference modeling to improve downstream performance. You will contribute to the development of core PyTorch-based training components and tooling that enable efficient experimentation and iteration. You will diagnose performance bottlenecks and drive improvements in training efficiency, data pipelines, and hardware utilization. You will document methodologies and results to support reproducibility and knowledge sharing across teams. This role is hybrid with a minimum of 2 in-office days per week in San Francisco.
Key facts
What you'll do
Design and implement scalable training and post-training pipelines for transformer-based language models across distributed compute environments.
Evaluate model quality, convergence behavior, and training efficiency to guide data, architecture, and training strategy decisions.
Partner with research teams to prototype and integrate novel training methods into production-grade workflows.
Optimize data ingestion, preprocessing, and caching systems to maximize GPU utilization and reduce iteration time.
Instrument and debug complex training runs to identify root causes of instability, degradation, or performance variance.
Collaborate with infrastructure engineers to improve cluster scheduling, resource isolation, and workload isolation.
Work with product and customer teams to adapt training and fine-tuning workflows for real-world deployment scenarios.
Define and maintain reliability and observability standards for training experiments and long-running model jobs.
Champion best practices in software engineering, versioning, and reproducibility for model training artifacts.
Support evaluation frameworks that assess model behavior across safety, robustness, and domain-specific benchmarks.
Enable developer workflows that simplify experiment configuration, monitoring, and comparison across runs.
Contribute to open-source components of the Lightning AI platform where appropriate to accelerate broader ecosystem adoption.
Drive automation of repetitive tasks in training lifecycle management to reduce manual overhead.
Mentor junior engineers by providing code review, design guidance, and technical leadership on model training challenges.
Engage with the latest research and translate promising ideas into robust, deployable systems.
Participate in on-call rotations to address critical training issues and ensure high availability of training platforms.
Requirements
Must have a Bachelor's, Master's, or PhD in Computer Science, Engineering, or a related quantitative field.
Must have 5+ years of experience building and training transformer-based language models or equivalent deep learning models.
Must have hands-on experience with large-scale distributed training frameworks and associated toolchains.
Must be proficient in PyTorch and Python for building and extending training and inference systems.
Must have strong software engineering practices including testing, debugging, and version control.
Must have experience optimizing training performance, including kernel awareness, memory efficiency, and compute utilization.
Must have experience working with production ML infrastructure and CI/CD patterns for model training.
Must have excellent communication skills to collaborate with researchers, engineers, and stakeholders across time zones.
Nice to have
Experience with post-training methods such as supervised fine-tuning, reinforcement learning from human feedback, and model merging.
Experience with efficient attention kernels, quantization, and model compression techniques.
Experience contributing to open-source machine learning frameworks.
Experience with multi-node and TPU training workflows.
Experience serving inference workloads and understanding deployment constraints.
Practical notes
This role is hybrid with a minimum of 2 in-office days per week in San Francisco.
Travel is not required for this role.
This role is eligible for US work authorization without sponsorship.
Candidates must be authorized to work in the United States now and for the duration of employment.
The listed compensation range reflects minimums and maximums set by Lightning AI for this role.