Senior Software Engineer, Robinhood Command Center
Job description
About the role
The Robinhood Command Center is a new reliability unit tasked with managing production incidents across the company. You will join a founding team responsible for building the processes and tools that detect and mitigate system issues. This role focuses on operational excellence and incident leadership rather than maintaining specific product services. You will define how the company responds to critical failures and drive improvements in system resilience. The position requires ownership of end-to-end reliability workflows from detection through resolution. You will partner closely with service owners to ensure consistent execution of mitigation strategies. This is a leadership position that shapes the operational maturity of the entire engineering organization. Your work will directly influence the stability and trustworthiness of Robinhood's platform.
Key facts
What you'll do
- Lead the long-term reliability and observability strategy for company infrastructure.
- Coordinate incident mitigation by guiding service owners through traffic shifts, rollbacks, and decision-making.
- Develop and maintain company-wide incident management procedures.
- Manage global dashboards and alerts focused on critical user journeys and business-impact metrics.
- Improve incident response metrics like MTTD and MTTR through tool development and process refinement.
- Oversee post-incident governance, including SEV reviews and postmortems.
- Design failure mitigation strategies to prevent full-region or full-datacenter outages.
- Build frameworks for monitoring and observability across hundreds of systems.
- Provide executive-level reporting on service quality and reliability.
- Mentor engineers and contribute to the growth of the engineering culture.
Requirements
- 5+ years of software engineering experience with production systems.
- 2+ years of experience in production operations, distributed systems, infrastructure, or reliability engineering.
- Hands-on experience as an incident commander, IMOC, or primary on-call responder.
- Proficiency in multi-cluster or multi-region architectures, failover strategies, and capacity planning.
- Deep understanding of fault-tolerant architecture and observability frameworks.
- Demonstrated success in improving availability, MTTR, MTTD, or customer impact metrics.
- Strong communication skills for cross-functional collaboration during high-pressure incidents.
- Experience working with financial services or regulated environments is expected.
Skills & tools
- Modern observability stacks such as Prometheus, OpenTelemetry, and Grafana.
- Proficiency with alerting systems and incident response platforms.
- Fluency in at least one general-purpose programming language for automation.
- Familiarity with cloud infrastructure and container orchestration platforms.
- Understanding of database replication, caching layers, and network fundamentals.
- Knowledge of runbooks, playbooks, and standard operating procedures.
- Exposure to log aggregation and distributed tracing tools.
- Ability to interpret service level objectives and service level indicators.
Nice to have
- Background working on high-availability systems that serve critical users.
- Experience with chaos engineering or failure injection practices.
- Contributions to open-source observability projects.
- Familiarity with regulatory compliance frameworks relevant to finance.
- Experience mentoring junior engineers through technical growth programs.
- Track record of driving organizational change in reliability practices.
Practical notes
- Compensation includes bonus opportunities, equity ownership, and 401(k) matching.
- Benefits include 100% paid employee health insurance, 90% dependent coverage, fertility benefits, mental health support, and a lifestyle wallet for wellness and learning.
- The office provides catered meals and events.
- Engagement is Full-time, in-office at least 3 days per week.
- This role is based in New York, NY.
- Candidates must be authorized to work in the United States without sponsorship for this position.
- The listed salary range applies to the New York, NY location.
The Senior Software Engineer, Robinhood Command Center role is positioned at the intersection of engineering leadership and operational resilience. You will be responsible for establishing the incident response culture across the company. This involves creating clear procedures that align technical teams during the most stressful situations. The position demands a balance of technical depth and soft skills to navigate complex stakeholder environments. You will translate raw system data into actionable insights for both technical and executive audiences. The role requires comfort with ambiguity and the ability to lead through influence rather than authority. You will continuously evaluate the effectiveness of existing tools and processes and propose improvements. Collaboration with product, infrastructure, and support teams is essential to maintain a unified view of system health. The work you do will establish the foundation for how Robinhood manages reliability for years to come. This is an opportunity to build a function from the ground up and leave a lasting impact on the organization. If you are passionate about operational excellence and thrive under pressure, this role is a chance to define the future of incident management at a major financial services company.