
Staff Software Engineer, Event Streaming Systems
Job description
USA.
About the role
DoorDash is seeking a Staff Software Engineer to define the technical roadmap and architecture for our event streaming infrastructure. This individual contributor role focuses on solving complex distributed systems challenges to support our global event-driven operations. You will own the design and execution of the multi-year vision for the event streaming platform, ensuring alignment with business needs and long-term scalability. In this position, you will lead critical initiatives such as migrating to object-store-backed, diskless streaming architectures while maintaining high reliability and performance. The role requires deep analysis of streaming internals to resolve intricate production issues and drive architectural improvements across the organization. You will act as a technical authority, influencing strategy and setting standards without relying on direct management control. Your work will directly impact the resilience and efficiency of the data pipelines that power DoorDash's core marketplace operations.
Key facts
What you'll do
- Establish the multi-year technical vision for the event streaming platform across the organization, defining clear goals and milestones.
- Lead high-impact initiatives, including the migration to object-store-backed, diskless streaming architectures that optimize cost and scalability.
- Perform deep-dive analysis into Kafka and modern streaming internals to resolve production issues and improve system reliability over time.
- Manage architectural decisions regarding replication, consensus, multi-region topology, and cost efficiency to balance performance and operational expenses.
- Provide technical mentorship and set engineering standards through design reviews, code contributions, and documentation efforts.
- Partner with leadership to evaluate build-versus-buy strategies, ensuring platform capabilities align with evolving business requirements.
- Drive the evaluation and adoption of emerging streaming technologies, frameworks, and tools to maintain competitive advantage.
- Collaborate with cross-functional teams to identify bottlenecks and implement robust solutions for high-throughput event processing.
- Own the reliability and operational excellence of streaming infrastructure, defining runbooks, alerting strategies, and incident response plans.
- Advocate for best practices in security, compliance, and data governance within the event streaming ecosystem.
- Explore and prototype new patterns in compute-storage separation to enhance flexibility and efficiency in large-scale deployments.
- Guide the selection and integration of third-party tools where appropriate, while maintaining alignment with open-source contributions.
Requirements
- 8+ years of professional experience building and operating large-scale distributed systems in production environments.
- Demonstrated expertise in Kafka, including log storage mechanisms, ISR models, metadata consensus (KRaft/ZooKeeper), and producer/consumer protocol behaviors.
- Proficiency in Java, Go, or similar languages with a strong focus on concurrency patterns and production-grade infrastructure development.
- Ability to lead cross-team initiatives and influence technical direction in the absence of direct authority, relying on persuasion and expertise.
- Strong understanding of Linux internals, networking protocols, storage systems, and JVM-based performance tuning techniques.
- Experience analyzing object-store-backed streaming trade-offs, such as compute-storage separation, data caching design, and latency implications.
- Proficiency using AI coding tools (e.g., Cursor, Claude Code, Codex) throughout the development lifecycle to enhance productivity and code quality.
- Track record of designing and implementing systems that meet stringent durability, consistency, and availability requirements.
- Experience working in fast-paced environments where ambiguity is common and technical leadership is essential.
- Willingness to engage with complex legacy systems while planning modernization efforts in a responsible and incremental manner.
Nice to have
- Hands-on experience with WarpStream, AutoMQ, or tiered storage implementations in Kafka environments.
- Contributions to open-source distributed systems projects such as Pulsar or Redpanda, demonstrating community engagement.
- Practical experience with Cruise Control, partition management, or broker tuning at large scale in demanding workloads.
- Background in designing multi-region platforms that emphasize high consistency, durability, and disaster recovery strategies.
- Proven track record of executing Tier-0 infrastructure migrations with zero customer impact and minimal operational disruption.
- Familiarity with monitoring and observability tools tailored to streaming platforms, including metrics collection and log analysis.
- Experience contributing to or maintaining internal libraries or frameworks used across engineering teams.
Skills & tools
Kafka, Java, Go, Linux, Distributed Systems, Object-store architectures, JVM tuning, AI coding assistants, version control, containerization, cloud infrastructure.
Practical notes
DoorDash uses Gem and Covey for recruitment and applicant evaluation. All candidates must go through these tools as part of the hiring process. The role offers comprehensive benefits, including 401(k) matching, 16 weeks of paid parental leave, medical, dental, and vision coverage, and flexible paid time off to support work-life balance. The position requires physical presence within the New York Metro Area, and candidates must be eligible to work in the United States without sponsorship for this role. There are no specific deadlines published for this position, and engagement is full-time. This position is not open to third-party agencies or consulting firms.