Software Engineer, Platform Infrastructure
Job description
About the role
Become a member of the infrastructure team responsible for constructing the foundational systems that enable distributed artificial intelligence development. This position focuses on engineering the control and data planes which empower users to seamlessly scale machine learning workloads from individual local machines up to massive cloud clusters. You will tackle complex challenges at the intersection of distributed systems, cloud infrastructure, and AI acceleration. The work directly impacts the reliability and performance of the platform used by developers and researchers globally. Collaboration with cross-functional teams is essential to deliver , production-grade infrastructure capabilities.
Key facts
What you'll do
- Design and scale services responsible for orchestrating Ray clusters across diverse cloud providers and on-premise data center environments.
- Optimize the control plane to efficiently handle massive distributed artificial intelligence and machine learning tasks at scale.
- Develop intelligent systems for resource scheduling and management within heterogeneous cluster environments containing varied hardware types.
- Elevate the performance, reliability, and observability of managed Ray workloads to meet stringent production standards.
- Manage the integration and lifecycle of hardware accelerators, specifically Graphics Processing Units and Tensor Processing Units.
- Own the container image lifecycle and dependency management strategy to ensure consistent, secure, and reproducible build artifacts.
- Actively participate in architectural design reviews and rigorous code quality assessments to maintain high engineering standards.
- Provide on-call rotation support to rapidly diagnose and resolve critical infrastructure incidents in coordination with field engineering teams.
- Implement automation frameworks for infrastructure provisioning, configuration management, and disaster recovery procedures.
- Collaborate with product teams to define technical requirements and translate them into scalable system designs.
- Conduct performance profiling and capacity planning to anticipate scaling bottlenecks before they impact users.
- Contribute to internal developer tooling and platforms that improve the velocity and safety of infrastructure changes.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical work experience demonstrating comparable depth.
- Minimum of three years of professional experience developing, deploying, and maintaining production-grade software systems.
- Demonstrated practical history of maintaining high-availability, scalable distributed systems in a production environment.
- Deep proficiency with cloud-native infrastructure patterns and services on at least one major provider: Amazon Web Services, Microsoft Azure, or Google Cloud Platform.
- Hands-on experience deploying, operating, and troubleshooting Kubernetes clusters in production settings.
- Strong grasp of cloud security principles, advanced networking concepts (VPC, CNI, load balancing), and authentication/authorization protocols (OAuth, OIDC, mTLS).
- Solid understanding of Linux operating system internals, including kernel scheduling, file systems (ext4, xfs), cgroups, namespaces, and containerization runtimes (containerd, runc).
- Fluency in Python and Go for systems programming, automation, and building backend services.
Nice to have
- Direct experience with the Ray framework for scaling AI and Python applications.
- Familiarity with observability stacks, specifically Prometheus for metrics collection and Grafana for visualization and alerting.
- Background in building or operating ML platform tooling, feature stores, or model serving infrastructure.
- Experience with Infrastructure as Code tools such as Terraform, Pulumi, or Crossplane.
- Knowledge of service mesh technologies like Istio or Linkerd for traffic management and security.
- Contributions to open-source projects in the cloud-native or distributed systems ecosystem.
Skills & tools
- Go
- Python
- Kubernetes
- AWS
- Azure
- GCP
- Linux
- Prometheus
- Grafana
- Ray
- Terraform
- Docker
- containerd
- gRPC
- Protocol Buffers
- CI/CD pipelines
Practical notes
Anyscale operates as an Equal Opportunity Employer and evaluates all applicants without regard to race, religion, color, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, genetic information, sexual orientation, gender identity or expression, or any other characteristic protected by applicable federal, state, or local laws. The role is based in Bengaluru, Karnataka, and requires the ability to work within Indian Standard Time business hours for effective collaboration with the global team. Visa sponsorship details are not explicitly listed in the source material; candidates requiring sponsorship should inquire directly during the recruitment process. Salary range information was not provided in the source listing.