Software Engineer (Ray Core)
Job description
About the role
The role advances Ray as a universal API for distributed applications through systems software work that defines the next generation of how compute scales. Engineers on the Ray Core team own the C++ backend that directly determines performance, reliability, and scalability for every Ray user in production. The team balances feature development, distributed libraries, test infrastructure, debugging, and architectural improvements to maintain a robust and forward-looking codebase. This position involves deep collaboration with cross-functional partners to translate complex requirements into efficient, production-grade systems software. The hire will be responsible for designing and implementing core subsystems that enable scalable machine learning workloads from laptop to cluster. Success in this role is measured by the stability, performance, and clarity of the distributed systems primitives provided to the community. The work directly impacts open source projects and influences how industries build and deploy AI applications at scale. Strong engineers combine technical depth with the ability to explain intricate systems decisions in plain, accessible words.
Key facts
What you'll do
Architectural improvements to Ray core are designed to simplify distributed programming and scale ML applications from laptop to cluster through careful systems abstraction.
High quality open source software is developed to lower the barrier for developers using Ray in production by ensuring robust interfaces and predictable behavior.
Cross-team projects are led to advance distributed libraries, runtime integration, and core subsystems while mentoring junior engineers to elevate the overall team capability.
Testing processes for Ray are improved to streamline releases and increase stability under large-scale workloads and stress conditions that mimic real production traffic.
Work on fault tolerance, performance, and reliability for distributed model training and inference, including tensor and pipeline parallelism to support next generation AI workloads.
Communication of technical work happens through talks, tutorials, and blog posts to reach a broader developer audience and foster a strong open source community.
Implementation of low level systems components focuses on C and C++ to ensure efficient memory management, concurrency, and responsiveness in demanding environments.
Collaboration with product and research teams helps align Ray core development with emerging use cases and evolving machine learning paradigms.
Requirements
Candidates must have at least 5 years of relevant work experience in distributed systems or backend engineering with a strong track record of delivering reliable systems.
Extensive experience building scalable and fault tolerant distributed systems is required for success in this role across heterogeneous environments.
Deep expertise in C and C++ and low level operating systems is necessary to contribute effectively to the Ray backend and its performance critical paths.
Solid background in algorithms, data structures, and system design is required to navigate the complexity of distributed computing challenges.
Knowledge of distributed model training and inference is preferred to better understand the workload characteristics and optimization opportunities.
Knowledge of GPU programming is preferred to engage with high performance computing paradigms and accelerate numerical workloads.
Practical experience with distributed testing, debugging, and profiling tools is expected to ensure quality and stability at scale.
The ability to write clear documentation and communicate technical tradeoffs supports collaboration across engineering organizations.
Practical notes
The role is based in San Francisco with standard U.S. work authorization requirements and adherence to E-Verify processes.
Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds where candidates solve a coding problem, discuss past projects, and answer system design questions.
Some loops include a take-home task that allows deeper exploration of engineering judgment and implementation style.
Final rounds typically cover team fit and give candidates a chance to ask questions to assess alignment with the team and company expectations.
Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer, so demonstrating structured thinking is crucial.
Practicing a few problems aloud and reviewing your own past projects are the best preparation to showcase your capabilities and learning agility.
Good to know
Ray is an open source project that supports scalable machine learning across many industries and diverse deployment scenarios.
The role focuses on systems programming in C++ for distributed computing infrastructure that underpins large scale ML applications.
Engineers work on core runtime, memory, and I/O subsystems that power large scale ML workloads across varied hardware and network conditions.
The team values clarity, reliability, and measurable improvements to stability and performance through rigorous experimentation and analysis.
Work often involves stress testing, debugging complex interactions between distributed components, and long term architectural planning to future proof the platform.
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels as demonstrated by expanding impact and technical leadership.
Some engineers move into management and lead teams of five to twenty people, while others remain on the technical track and deepen their expertise in core systems.
Growth follows demonstrated impact, not tenure alone, and is reflected in increasingly complex problem solving and cross team collaboration.
A typical engineering ladder has clear levels with defined expectations for scope, quality, mentorship, and ownership of critical deliverables.
Moving up usually requires owning outcomes end to end rather than completing assigned tickets, ensuring that solutions are robust, maintainable, and aligned with user needs.
Compensation details are not specified in the source material, so specific pay information is not included on this page. The role is full time based in San Francisco and requires standard U.S. work authorization, with the company operating under E-Verify to confirm employment eligibility. The position demands a Bachelor's degree or equivalent experience, along with at least five years of hands on experience in distributed systems or backend engineering. Applicants should be prepared to demonstrate deep expertise in C and C++ along with low level operating systems concepts, as these form the foundation for contributing effectively to the Ray Core team. The focus on architectural improvements, high quality open source software, and rigorous testing processes underscores the expectation for engineers who are meticulous, communicative, and driven by measurable impact.