Senior Site Reliability Engineer, Platform Infrastructure
Job description
About the role
This role designs, builds, and operates infrastructure for distributed AI workloads on the Anyscale cloud platform. The position integrates open-source Ray with proprietary products and ensures scalable, secure, and execution.
Data roles turn raw information into decisions. Analysts query databases and build dashboards. Data scientists build models that predict outcomes. Data engineers build the pipelines that move and store data. All three work closely with business teams and need a mix of statistics, coding, and communication. Nearly every modern company runs on data teams, from startups to banks. A strong portfolio of past analyses matters more than degrees in many hiring decisions.
Key facts
What you'll do
Control plane components are tuned to orchestrate cluster management, scheduling, and user access for large-scale, distributed AI/ML workloads.
Accelerator integration for GPUs and TPUs is optimized to satisfy demanding AI workload requirements.
Container image management and dependency resolution are implemented to ensure consistent and reliable deployments.
Requirements
A Bachelor's degree in Computer Science, Engineering, or equivalent practical experience is required as proof of foundational knowledge.
At least 3+ years of experience writing high-quality production code is required to demonstrate maturity and reliability in software delivery.
Hands-on experience building and maintaining highly available, scalable, and performant distributed systems in production environments is required.
Expertise in cloud-native technologies across AWS, Azure, and GCP and Kubernetes-based deployments for multi-cloud fluency is required.
Deep understanding of networking, security, and authentication mechanisms within cloud environments is required to protect and isolate workloads.
Familiarity with observability stacks including Prometheus and Grafana is required to monitor, debug, and optimize infrastructure.
Proficiency in Go and Python is required to implement services, operators, and integrations effectively.
Knowledge of low-level operating system foundations such as the Linux kernel, file systems, and containers is required to tune performance and reliability.
Practical notes
The role is based under US E-Verify participation and requires a Bachelor's degree as stated.
Typical interview steps
Data interviews commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies give a take-home analysis. Expect questions about past projects and the business impact of your work. Interviewers often evaluate how you communicate uncertainty and business impact, not only the math. Bringing a clean write-up of a past analysis to the interview is well received.
Good to know
Site reliability engineering in AI infrastructure combines cloud operations, distributed systems, and networking concepts.
Common tools in this field include Kubernetes, container orchestration frameworks, and observability stacks for monitoring and alerting.
Distributed AI workloads involve accelerator integration, large-scale scheduling, and data plane optimization.
Control plane and data plane responsibilities split operational concerns between orchestration and execution.
Open-source contributions and integration with projects like Ray are part of the day-to-day work.
Questions to ask
Useful questions for the interview: what a typical week looks like, how work is assigned, what tools the team uses, and how feedback works. Asking how the role has changed recently and what the team wishes it had known when joining is also reasonable. Questions about the manager's priorities are especially valued.
Career growth
Data careers grow toward senior analyst, staff data scientist, or data engineering lead. Many professionals specialize in machine learning, analytics, or infrastructure. Cross-functional work with product and engineering teams becomes more important at senior levels. The field changes quickly, so continuous learning is part of the job. Professionals who can translate numbers into decisions tend to advance fastest.