
Machine Learning Performance Engineer
Job description
About the role
Applied Intuition is seeking a Machine Learning Performance Engineer to own the critical work of maximizing the efficiency and throughput of large-scale machine learning infrastructure in our datacenter environment. The hire will be responsible for ensuring that distributed training runs and massive offline inference sweeps operate at peak cluster goodput while minimizing wasted compute resources. This role owns the end-to-end performance stack, from data loading and kernel execution to communication patterns and checkpointing strategies across many nodes. You will establish rigorous performance models and roofline analyses to expose where time and money are lost in our workloads. Close the gap between the theoretical capabilities of our accelerators and the actual achieved throughput by identifying and eliminating systemic bottlenecks. Collaboration with cross-functional teams is central, as you will partner with infrastructure, framework, and data teams to drive joint wins. Ultimately, your work will directly accelerate how quickly the company can iterate on models and process autonomy data by making every GPU hour count.
Key facts
What you'll do
- Profile and optimize distributed training end to end, including data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing for maximum cluster efficiency.
- Optimize large-scale offline and batch inference over petabyte-scale sensor logs by designing intelligent batching and scheduling strategies, applying quantization and low-precision execution, and improving accelerator saturation during long-running sweeps.
- Establish roofline and performance models for our workloads, quantify the gap between achieved and theoretical performance, and stack-rank optimization opportunities by both impact and implementation effort.
- Improve multi-node scaling efficiency through refined sharding and parallelism strategies, collective communication optimization, and maximized interconnect utilization while addressing memory-bandwidth and kernel-fusion bottlenecks.
- Drive cluster goodput by systematically reducing GPU idle time caused by input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery in long-running jobs.
- Build and maintain benchmarking, observability, and regression-detection tooling that prevents performance degradation as models, data, and code evolve over time.
- Collaborate with engineers across functions to solve complex data and compute problems at scale, ensuring that cross-layer optimizations are aligned with real workload requirements.
- Contribute to a team culture that values effective collaboration, technical excellence, and continuous innovation in performance engineering practices.
- Analyze real-world autonomy log processing pipelines to uncover inefficiencies and design targeted improvements that reduce data processing latency and cost.
- Partner with infrastructure teams to ensure that scheduling, network topology, and storage configurations are tuned for optimal ML throughput.
- Evaluate emerging accelerator features and low-level kernel optimizations to determine applicability to current workload patterns.
- Translate performance insights into clear documentation and actionable recommendations for both engineers and stakeholders.
- Lead focused investigations on specific performance regressions, using data-driven methods to isolate root causes and validate fixes.
- Work closely with product teams to balance performance objectives with delivery timelines and resource constraints on critical initiatives.
Requirements
- Hands-on ML performance engineering experience, including profiling, roofline analysis, throughput optimization, and root-cause investigation in production systems at scale.
- Proven experience with distributed multi-node training at scale using frameworks such as FSDP, DeepSpeed, or Megatron, and with collective communication libraries like NCCL.
- Deep familiarity with GPU or accelerator performance concepts, including memory bandwidth, kernel launch overhead, occupancy, quantization strategies, and collective communication patterns.
- Extensive experience with high-throughput or batch inference systems such as NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, or Ray in production environments.
- Fluency in Python and proficiency in at least one systems-level language such as C++ is required for low-level optimization and instrumentation.
- Excellent debugging, analytical, and problem-solving skills, with the ability to dissect complex performance issues in large distributed systems.
- A deep understanding of machine learning foundations, including training and inference workloads, data pipelines, and model architectures.
- The ability to develop technical solutions for problems with no established precedent, relying on first-principles reasoning and empirical measurement.
Practical notes
This is an in-office role, and the expectation is that employees work primarily from their Applied Intuition office 5 days a week. However, flexibility is recognized, allowing for occasional remote work, such as starting the day with morning meetings from home before heading to the office or adjusting schedules to accommodate family commitments when feasible.