
Software Engineer, Spark Platform
Job description
USA.
About the role
Building the Platform That Powers Data at Scale
The Spark Platform team is responsible for the infrastructure that powers DoorDash's data, analytics, and machine learning. We manage the entire Apache Spark ecosystem, including the execution runtime, remote shuffle service, cluster scheduling, and the reliability tooling that ensures stability. Our work operates at significant scale, supporting an expanding set of workloads and consumer teams. Tackling the complexity of orchestrating thousands of deployments requires deep investment in runtime optimization, systems architecture, and resilient tooling. You will be hands-on across the entire lifecycle of our in-house Spark deployment, which serves the entire company. Your focus will span runtime performance, scheduler efficiency, cluster lifecycle automation, and the observability that keeps the platform reliable. You will tackle high-use problems wherever they appear, moving fluidly between layers of the stack. Close collaboration with both the core team and platform consumers will ensure solutions meet real-world needs.
Specific responsibilities include designing and implementing observability and incident automation to sustain on-call health as scale increases. You will operate Kubernetes controllers and custom resources to automate cluster provisioning, upgrades, and recovery at a level where manual processes no longer suffice. You will drive multi-tenant scheduling and executor placement strategies that maximize resource efficiency for many teams. Building and refining the remote shuffle service and runtime upgrade pathways will be essential to support growing analytics demands. You will analyze performance regressions, define incremental improvements, and validate changes through measurement. Reliability enhancements will be delivered through controlled rollout strategies, and you will partner on the architecture of shuffle and scheduling components. This is a hybrid role requiring relocation to San Francisco, Sunnyvale, Seattle, or New York City.
Key facts
What you'll do
- Execute the full scope of work for the Software Engineer, Spark Platform role at DoorDash, translating requirements into robust technical implementations.
- Adhere strictly to the defined scope requirements for the Software Engineer, Spark Platform position, ensuring all deliverables align with team objectives.
- Satisfy the practical bars outlined in the hiring criteria, demonstrating consistent eligibility throughout the review process.
- Architect and implement observability frameworks and incident automation systems to maintain on-call health as platform scale increases exponentially.
- Operate Kubernetes controllers and custom resources to automate cluster lifecycle events, including provisioning, upgrades, and recovery workflows.
- Drive strategic multi-tenant scheduling and executor placement initiatives to maximize resource efficiency across numerous consumer teams.
- Develop and refine the remote shuffle service and runtime upgrade pathways to accommodate surging analytics demands and evolving workloads.
- Analyze performance regressions in depth, define targeted incremental improvements, and validate changes through rigorous measurement and experimentation.
- Deliver reliability enhancements using controlled rollout strategies that minimize risk and ensure stable platform operations.
- Partner closely with architecture teams on the design of shuffle and scheduling components to ensure long-term scalability and resilience.
- Collaborate continuously with the core Spark team and platform consumers to ensure solutions address real-world operational needs effectively.
- Balance innovation with operational pragmatism, prioritizing solutions that reduce toil and improve maintainability at scale.
- Contribute to end-to-end ownership of platform features, from initial design through deployment and post-launch optimization.
- Leverage deep systems thinking to troubleshoot complex issues that span multiple layers of the distributed stack.
- Champion best practices in infrastructure as code, enabling reproducible and auditable platform management.
Requirements
- Hold a Bachelor of Science, Master of Science, or Doctor of Philosophy in Computer Science or an equivalent technical background.
- Bring extensive experience operating production-grade distributed systems within cloud-native environments.
- Demonstrate a strong background in Apache Spark platform operations, with emphasis on runtime management, shuffle infrastructure, and scheduling.
- Possess hands-on expertise with Kubernetes, including controllers, operators, custom resources, and an understanding of failure modes in multi-tenant clusters.
- Show familiarity with batch-oriented or big-data schedulers, and experience with the Spark-on-Kubernetes operator where applicable.
- Be fluent in defining observability strategies, including SLOs, SLIs, and the use of monitoring stacks such as Prometheus and OpenTelemetry.
- Apply practical cloud experience, particularly with AWS services, VPC networking, instance lifecycle management, and autoscaling primitives.
- Write professional-grade code in Python, Go, Scala, or Java, and maintain proficiency in SQL for data-intensive operations.
- Favor incremental delivery, rigorous measurement, and the reduction of operational toil in all platform work.
- Exhibit strong problem-solving skills and the ability to debug complex issues in distributed systems at scale.
- Communicate effectively with both technical and non-technical stakeholders to align on priorities and trade-offs.
- Thrive in a fast-paced environment where ownership and initiative are essential for success.
- Commit to following defined processes and contributing to continuous improvement of platform practices.
- Maintain reliability and stability focus when implementing new features and infrastructure changes.
Nice to have
- Only if the SOURCE document specifies preferred qualifications, list them here. The SOURCE does not specify additional preferred qualifications beyond the stated requirements.
Practical notes
- Your application will be reviewed based on the information provided. Confirm all details on the official application page before submitting.
- This is a hybrid role requiring relocation to San Francisco, Sunnyvale, Seattle, or New York City.