Senior Software Engineer
Job description
About the role
Roblox's Cache team is seeking a Senior Software Engineer to architect and operate next-generation caching infrastructure that serves tens of millions of daily users. This position focuses on building a revolutionary caching solution targeting sub-millisecond latency, horizontal scalability, and dramatic cost reduction while supporting the company's ambitious goal of reaching one billion daily active users. You'll work on large-scale distributed systems within the Infrastructure Storage organization, reporting to the Cache team's Engineering Manager. The role combines deep technical innovation with operational excellence, requiring expertise in distributed systems and a passion for solving complex infrastructure challenges at unprecedented scale.
Key facts
What you'll do
- Spearhead the architectural evolution toward a next-generation multitenant caching service built on ValKey, implementing comprehensive isolation mechanisms for data, resources, and failure domains across all tenant environments
- Execute systemic performance optimizations addressing critical challenges including head-of-line blocking scenarios, hot key management strategies, and maximizing CPU and memory resource utilization across distributed physical machine clusters
- Architect and implement comprehensive automation frameworks covering development workflows, chaos engineering practices with fault and latency injection capabilities, and continuous monitoring systems for mission-critical services operating around the clock
- Target and maintain exceptional service reliability metrics of 99.99% or higher availability while ensuring elastic scalability to accommodate dynamic workload patterns and traffic fluctuations
- Lead technical design review sessions with cross-functional stakeholders, evaluating architectural proposals and ensuring alignment with organizational standards and scalability requirements
- Conduct rigorous performance benchmarking exercises to identify optimization opportunities and validate system behavior under various load conditions and failure scenarios
- Facilitate and participate in failure drill exercises to test system resilience, validate recovery procedures, and identify potential weaknesses before they impact production environments
- Drive blameless post-incident retrospective meetings, fostering a learning culture that transforms operational incidents into opportunities for systemic improvement and knowledge sharing
- Provide technical mentorship to engineering team members, elevating their expertise in distributed systems, caching technologies, and infrastructure best practices
- Cultivate deep domain knowledge across the organization by facilitating knowledge transfer sessions and documentation efforts spanning Storage, Platform, and Product engineering teams
- Collaborate on reducing cluster lifecycle management overhead from hours to seconds, eliminating operational burden for service owners and enabling rapid capacity expansion
- Contribute to the team's mission of achieving 90% cost reduction while scaling infrastructure to support exponentially growing user populations and usage patterns
Requirements
- Bachelor's degree in Computer Science or equivalent professional experience demonstrating comparable technical foundation and problem-solving capabilities
- Minimum of six years of hands-on software engineering experience with demonstrated progression in technical complexity and responsibility
- Extensive expertise in designing, building, and operating large-scale distributed systems with proven understanding of consistency models, partition tolerance, and failure handling
- Strong infrastructure engineering background with practical experience running Active/Active distributed systems on container orchestration platforms such as Kubernetes or Nomad
- Demonstrated proficiency in Go programming language with ability to write production-grade, maintainable code for high-performance systems
- Solid experience with C++ for performance-critical components and understanding of memory management, concurrency patterns, and systems programming
- Proven track record resolving massive-scale performance bottlenecks and architectural limitations in distributed environments
- Practical experience addressing challenges such as decentralized Gossip protocol limitations or mitigating partial failure scenarios in distributed architectures
- Hands-on expertise with modern observability and telemetry stacks including Prometheus for metrics collection, Grafana for visualization, AlertManager for alerting, and Kibana for log analysis
- Strong builder mindset with ability to translate architectural vision into concrete implementation plans and operational systems
Nice to have
- Active contributions to or maintenance responsibilities for major open-source caching projects such as Redis, ValKey, or Memcached with demonstrable community involvement
- Advanced experience extending cache functionality through custom Redis modules written in C or Rust, or complex Lua scripting for specialized caching behaviors
- Deep expertise in memory allocator tuning and optimization, particularly with jemalloc configuration for cache workloads
- Practical experience implementing or operating caching proxy solutions such as Twemproxy or Envoy Redis filter in production environments
- Background designing and deploying complex multi-tiered caching architectures with various topology patterns optimized for specific access patterns
- Experience with chaos engineering methodologies and tools for validating distributed system resilience under adverse conditions
Skills & tools
Go, C++, ValKey, Redis, Memcached, Kubernetes, Nomad, Prometheus, Grafana, AlertManager, Kibana, distributed systems architecture, container orchestration, chaos engineering, performance optimization, memory management, observability platforms, caching proxies, Lua scripting, jemalloc
Practical notes
- This position is based at Roblox headquarters in San Mateo, California, requiring physical presence in the office three days per week
- The compensation package includes base salary within the stated range, equity compensation, and comprehensive benefits as detailed on the company's benefits page
- Actual base salary determination considers professional background, training, work experience, business needs, and market demand, potentially varying from the stated range
- Roblox maintains an equal employment opportunity policy prohibiting discrimination based on race, color, religion, age, sex, national origin, disability status, genetics, veteran status, sexual orientation, gender identity or expression, or other protected characteristics
- Reasonable accommodations are available for candidates with qualifying disabilities or religious beliefs throughout the recruiting process
- The company cannot employ candidates with work authorization related to certain U.S. visa categories and does not provide H-1B sponsorship for this role at this time
- The Cache team recently published engineering blog posts detailing their technical work, available for review to understand team initiatives and technical approach