Senior Software Engineer, Machine Learning Infrastructure
Job description
Senior Software Engineer, Machine Learning Infrastructure at DoorDash USA.
About the role
This role involves designing and implementing production infrastructure for Generative AI. You will lead the technical direction for open-weight model platforms, focusing on inference and fine-tuning. This position is for a senior engineer who thrives on solving complex, high-impact system challenges.
Key facts
What you'll do
- Lead the architectural design of infrastructure that enables DoorDash teams to deploy Generative AI solutions from concept to production.
- Manage and enhance the open-weight model serving stack, including real-time GPU endpoints, high-throughput batch inference, and fine-tuning methods.
- Develop scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning to support customer and internal automation.
- Improve the cost-efficiency and reduce the latency of GPU inference, converting lengthy batch jobs into much shorter durations and significantly cutting inference expenses.
- Create platforms that facilitate rapid experimentation while adhering to production standards for performance, monitoring, and operational stability.
- Collaborate with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to integrate new GenAI capabilities into core platform features.
- Define the technical roadmap for DoorDash's central GenAI platform, exploring areas like reinforcement learning and agent optimization to enable future AI products.
Requirements
- Bachelor's, Master's, or PhD in Computer Science or a related field.
- 6+ years of professional experience in software engineering.
- Strong understanding of backend engineering principles, particularly in Python and distributed systems.
- Proven experience in designing and managing production services, APIs, data pipelines, or ML infrastructure at scale.
- Practical experience operating systems in a production environment, covering observability, debugging, reliability, incident response, and performance/cost optimization.
- Extensive hands-on experience with LLM inference or fine-tuning of open-weight models in production, including serving aspects like latency, throughput, batching, autoscaling, GPU utilization, or fine-tuning methods such as SFT, DPO, or LoRA.
- Demonstrated technical leadership, including leading design efforts in complex, evolving technical domains, mentoring other engineers, and transforming user needs into reusable platform capabilities.
- Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) throughout the software development lifecycle, from design and code generation to testing, monitoring, and release.
Nice to have
- Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in a production setting.
- Experience with distributed or multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluation.
- Background in GPU performance optimization, such as multi-node/distributed inference, KV-cache/memory optimization, quantization (FP8/INT8/AWQ/GPTQ), or cold-start/throughput tuning.
- Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (e.g., Modal), or high-throughput batch systems.
- Experience with LLM gateways, model routing, vendor abstraction, or cost attribution.
- Experience building developer platforms, internal platforms, or self-serve infrastructure.
- Experience building and deploying AI agents or MCP servers in production.
- Experience with evaluation systems, LLM observability, tracing, RAG, search, or vector databases.
Skills & tools
- Python
- Distributed Systems
- GPU serving
- LLM inference
- LLM fine-tuning
- SFT/DPO/LoRA
- APIs
- Data pipelines
- Observability
- AI coding tools (e.g., Claude Code, Codex, Cursor)
Practical notes
Compensation:
I4: $137,100 - $201,600 USD
I5: $167,800 - $246,800 USD
I6: $203,500 - $299,300 USD
Base salary is localized by work location. This role also includes equity grants.
Benefits include a 401(k) plan with employer matching, 16 weeks of paid parental leave, wellness benefits, commuter benefits match, flexible paid time off, 80 hours of paid sick time per year, medical, dental, and vision benefits, 11 paid holidays, disability and basic life insurance, family-forming assistance, and a mental health program. DoorDash uses an automated recruitment tool called Gem, which assists in evaluating job qualifications. This tool supports human decision-making and is subject to human review. Data collected is retained according to the Candidate Privacy Policy. A bias audit summary for Gem is publicly available on the Careers Page.