Software Engineer, Ray Data
Job description
About the role
Anyscale is seeking a seasoned engineer to advance the Ray Datasets library, a foundational component for scalable machine learning pipelines used by major technology firms. This position centers on optimizing high-performance data processing built atop Apache Arrow and the Ray Core C++ backend. You will tackle complex challenges in distributed systems, focusing on scheduler interactions, memory management, and I/O subsystems. The role offers the opportunity to shape the future of streaming workloads integration, such as Beam on Ray, within a widely adopted open-source ecosystem. Collaboration with adjacent ML libraries like Train, RLlib, and Serve is a daily aspect of the workflow.
Key facts
What you'll do
- Architect and implement high-quality open-source software for the Ray project to lower the barrier for distributed programming across diverse hardware environments.
- Identify performance bottlenecks in Ray Datasets at massive scale and devise architectural improvements leveraging Apache Arrow primitives and the Ray object manager.
- Design and build robust integration pathways between Ray Data and popular ML training frameworks as well as heterogeneous data sources.
- Develop comprehensive stability and stress testing infrastructure to ensure smooth, reliable release cycles for the Ray ecosystem.
- Lead strategic initiatives to incorporate streaming workloads into the platform, including efforts related to Beam on Ray integration.
- Differentiate data operations specifically for the Anyscale hosted Ray service to provide enhanced value over the open-source offering.
- Collaborate intimately with Ray Core engineers on scheduler logic, memory subsystems, and I/O pathways to optimize end-to-end data flow.
- Partner with teams maintaining Ray Train, RLlib, and Serve to ensure seamless interoperability and unified user experiences.
- Conduct rigorous evaluation of system design trade-offs involving fault tolerance, scalability, and latency in distributed data processing.
- Author technical documentation, blog posts, and presentations to communicate complex engineering concepts to the broader developer community.
- Mentor peers on best practices for distributed systems development, code review standards, and performance profiling methodologies.
- Participate in on-call rotations or incident response for critical production issues affecting Ray Data stability in customer environments.
Requirements
- Minimum of five years of professional software engineering experience in relevant domains.
- Deep theoretical and practical command of algorithms, data structures, and large-scale system design principles.
- Proven track record of building, operating, and debugging scalable, fault-tolerant distributed systems in production settings.
- Hands-on experience with data processing engines, database internals, or frameworks such as Apache Spark or Dask.
- Strong proficiency in Python and C++ given the library's reliance on Arrow and the Ray Core backend.
- Demonstrated ability to navigate and contribute to large, complex open-source codebases with multiple stakeholders.
- Experience optimizing data layout, serialization, and memory management for analytical workloads.
- Familiarity with Kubernetes or container orchestration platforms for deploying distributed applications.
Nice to have
- Prior contributions to the Ray project, Apache Arrow, or similar open-source distributed computing frameworks.
- Experience with streaming processing systems such as Apache Beam, Apache Flink, or Kafka Streams.
- Background in machine learning infrastructure, feature stores, or ML pipeline orchestration tools.
- Knowledge of cloud provider managed services (AWS, GCP, Azure) for data analytics and storage.
- Experience presenting at technical conferences or authoring widely read technical blog posts.
- Understanding of GPU-accelerated data processing and heterogeneous compute scheduling.
Skills & tools
Python, C++, Apache Arrow, Ray Core, Distributed Systems, Apache Spark, Dask, Apache Beam, Kubernetes, System Design, Performance Optimization, Open Source Development, ML Infrastructure, Data Engineering.
Practical notes
The role is based in Bengaluru, Karnataka, and operates as a full-time engagement within the Engineering department. Anyscale has raised over $250 million from investors including Andreessen Horowitz, NEA, and Addition, backing the commercialization of the Ray open-source project. The technology stack is trusted by industry leaders such as OpenAI, Uber, Spotify, Instacart, and Cruise. Visa sponsorship details are not explicitly mentioned in the source material; candidates should inquire directly during the recruitment process. The compensation structure was not disclosed in the listing. Anyscale is an Equal Opportunity Employer committed to evaluating candidates without regard to protected characteristics.