Software Engineer, Ray Core
Job description
About the role
Anyscale is seeking a backend engineer to join the Ray Core team, which owns the C++ runtime powering the open-source Ray framework. This position focuses on enhancing the performance, reliability, and scalability of the distributed scheduler, memory management, and language runtime integration layers. You will tackle complex distributed systems challenges to ensure Ray can efficiently support massive AI workloads, from training to inference. The role involves a mix of architectural improvements, feature development, and hardening the testing infrastructure to guarantee smooth releases. Collaboration is key, as you will help define the primitives that higher-level ML libraries and end-users rely on daily.
Key facts
What you'll do
- Design and implement high-performance C++ components for the Ray Core backend, including the distributed scheduler, object store, and runtime environments.
- Diagnose and resolve bottlenecks in large-scale distributed workloads to improve throughput and reduce latency for machine learning applications.
- Architect and build fault-tolerance mechanisms, such as high-availability features for the head node and automatic recovery protocols for worker failures.
- Develop and maintain comprehensive stress testing and stability infrastructure to catch regressions before they reach production releases.
- Collaborate with the open-source community and internal stakeholders to review pull requests, triage issues, and guide the technical roadmap for core subsystems.
- Optimize memory management and I/O subsystems to efficiently handle heterogeneous hardware resources, including GPUs and TPUs.
- Integrate support for advanced distributed training paradigms, such as tensor parallelism and pipeline parallelism, directly into the runtime layer.
- Author technical documentation, design documents, and blog posts to communicate complex architectural changes to the broader engineering community.
- Mentor junior engineers on distributed systems best practices, debugging techniques, and performance profiling methodologies.
- Participate in on-call rotations to ensure the reliability of critical Ray Core services and respond to production incidents.
- Evaluate and adopt new systems technologies or algorithms that can enhance the scalability of the control plane and data plane.
- Contribute to the Ray project's governance processes, including RFC reviews and release management activities.
Requirements
- Minimum of two years of professional software engineering experience building production-grade systems.
- Deep expertise in C++ (modern standards preferred) with a strong grasp of concurrency, memory management, and low-level optimization.
- Proven track record designing, implementing, and operating scalable, fault-tolerant distributed systems in a Linux environment.
- Solid foundation in computer science fundamentals, including algorithms, data structures, operating systems, and network protocols.
- Experience debugging complex multi-threaded and multi-process applications using tools like gdb, perf, sanitizers, or eBPF.
- Familiarity with containerization technologies (Docker, Kubernetes) and cloud infrastructure (AWS, GCP, or Azure).
- Ability to write clear, maintainable, and well-tested code with a focus on long-term reliability.
- Strong verbal and written communication skills for collaborating across time zones and engaging with the open-source community.
Nice to have
- Direct experience with distributed machine learning training and inference frameworks, specifically tensor parallelism and pipeline parallelism techniques.
- Background in GPU programming using CUDA, HIP, or related accelerator APIs for high-performance computing.
- Prior contributions to the Ray project or other major open-source distributed systems (e.g., Apache Spark, Flink, Kubernetes).
- Knowledge of Ray's architecture, including the Global Control Service, Object Store, and Raylet components.
- Experience building developer tools, SDKs, or runtime systems for data science and ML practitioners.
- Familiarity with formal verification, model checking, or advanced testing methodologies like chaos engineering.
Skills & tools
C++, Python, Ray, Distributed Systems, Linux Kernel, gRPC, Protocol Buffers, Bazel, Git, GitHub Actions, Kubernetes, Docker, AWS, GCP, Azure, CUDA, Tensor Parallelism, Pipeline Parallelism, gdb, perf, AddressSanitizer, ThreadSanitizer, eBPF, Prometheus, Grafana.
Practical notes
This is a full-time position based in Bengaluru, Karnataka. The role requires close collaboration with a globally distributed team, necessitating flexibility for occasional meetings across US and European time zones. As a core maintainer of an open-source project, you are expected to engage publicly on GitHub, participate in design discussions, and contribute to the project's long-term technical direction. Anyscale offers a competitive compensation package including equity, comprehensive health benefits, and a stipend for home office setup or co-working space access. The company sponsors work visas for eligible candidates relocating to or currently residing in India. The interview process typically involves a recruiter screen, a coding assessment focused on systems problems, a system design interview, and a behavioral round with the hiring manager and team members.
META
Company: Anyscale
Title: Software Engineer, Ray Core
Listed
location: Bengaluru, Karnataka
Job type: full_time
Department: Engineering
Employment: FullTime
Compensation JSON: {"compensationTierSummary":null,"scrapeableCompensationSalarySummary":null,"compensationTiers":[],"summaryComponents":[]}
About the company
At Anyscale https://www. Anyscale. com/, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels.