Research Engineer
Job description
About the role
Lightning AI is seeking a Research Engineer to improve the performance, reliability, and scalability of AI systems on our platform. You will own the design and execution of experiments that validate architectural choices for large-scale AI workloads. This role requires you to translate novel research ideas into robust, production-grade components that operate reliably at scale. You will be responsible for diagnosing systemic inefficiencies and proposing concrete engineering solutions that impact the entire product stack. The ideal candidate thrives on turning ambiguous problems into measurable improvements in speed, stability, and resource utilization. You will act as a technical bridge between cutting-edge research prototypes and the high-availability infrastructure serving enterprise customers. Your work will directly influence how the Lightning platform scales to meet the demands of the most challenging AI workloads.
Key facts
What you'll do
- Oversee the optimization of large-scale training and inference performance across distributed systems and heterogeneous accelerators.
- Conduct deep analysis alongside customers to pinpoint workload bottlenecks and architect resilient production AI systems.
- Architect, build, and maintain high-throughput inference pipelines, scalable model serving systems, and performance-centric engineering tools.
- Design and deliver observability, profiling, and debugging utilities that provide deep insight into model execution and system behavior.
- Integrate performance enhancements directly into the Lightning ecosystem through clean, maintainable APIs and automated workflows.
- Work closely with hardware partners from NVIDIA, TPU, and other accelerator ecosystems to co-develop efficient execution strategies.
- Drive the creation of open-source contributions and high-quality technical documentation that clarify system capabilities and limitations.
- Evaluate and prototype inference optimization methods such as quantization, speculative decoding, mixed precision, and memory-efficient training techniques.
- Implement low-level kernels and integrations using CUDA, Triton, TensorRT, vLLM, SGLang, or Dynamo to unlock maximum hardware throughput.
- Champion best practices in software engineering by focusing on API consistency, debuggability, and reliable production code delivery.
Requirements
- Demonstrate advanced proficiency with deep learning frameworks, with a core focus on PyTorch and its ecosystem.
- Bring hands-on experience managing large-scale training or inference workloads in demanding production environments.
- Show a firm grasp of distributed systems concepts, including data, model, and pipeline parallelism, as well as elastic scaling and checkpointing strategies.
- Exhibit strong software engineering capabilities in API design, debugging complex systems, and delivering production-grade code.
- Prove a track record of identifying and resolving performance bottlenecks within machine learning infrastructure stacks.
- Hold a Bachelor degree in Computer Science, Engineering, or a closely related technical field.
- Display the ability to operate effectively in a fast-paced, cross-functional environment with shifting priorities and tight deadlines.
- Commit to working in a hybrid model with a minimum of two days in the office per week across London, New York, San Francisco, or Seattle locations.
- Meet compensation bands that align with local market standards and internal equity considerations for the listed regions.
- Adhere to company policies regarding employment eligibility and legal work authorization in your country of residence.
Nice to have
- Hands-on familiarity with inference optimization methods like quantization, speculative decoding, mixed precision, or memory-efficient training.
- Direct experience with performance-critical frameworks and tools such as CUDA, Triton, TensorRT, vLLM, SGLang, or Dynamo.
- A history of meaningful contributions to open-source machine learning or infrastructure projects with public repositories.
- Previous experience within startups or highly collaborative technical environments where impact was measured by execution speed.
- Possession of an advanced degree, such as a Master or PhD, in artificial intelligence, machine learning, or computer systems.
Practical notes
This is a full-time position with total compensation that includes a discretionary bonus and equity components. The salary range specified for this role is $120,000 to $250,000 USD base salary, adjusted for geographic location and individual qualifications. Benefits coverage includes medical, dental, and vision insurance for eligible employees. Retirement savings options include 401(k) matching for United States-based staff or pension contributions for United Kingdom-based staff. You will enjoy unlimited PTO, a two-week winter break, paid parental leave, and a four-week paid sabbatical after four years of service. Additional support is provided through in-office meals, a professional development allowance, wellness stipends, and WFH equipment stipends. Travel requirements are minimal, and the role operates under a hybrid schedule that requires two days of in-office presence per week. Employment is contingent on successful completion of any applicable visa authorization or work eligibility checks where required.