Distributed LLM Inference Engineer
Job description
About the role
Distributed LLM Inference Engineers deliver performance breakthroughs for large-scale inference and help Anyscale lead AI infrastructure. This role defines the core technical foundation for scaling language models in production.
Data roles turn raw information into decisions. Analysts query databases and build dashboards. Data scientists build models that predict outcomes. Data engineers build the pipelines that move and store data. All three work closely with business teams and need a mix of statistics, coding, and communication. Nearly every modern company runs on data teams, from startups to banks. A strong portfolio of past analyses matters more than degrees in many hiring decisions.
Key facts
What you'll do
High-throughput, low-latency inference systems are built and iterated quickly with product teams to serve open-source Ray users and enterprise customers. These solutions span batch and online workflows to meet scalable AI infrastructure goals.
Ray Data and the LLM engine are integrated to achieve low-cost inference across the entire stack. The integration connects data processing with model execution for efficient large-scale workloads.
Open-source components such as vLLM are integrated into Anyscale solutions. Engineers also contribute extensions and improvements back to the community to advance shared tooling.
The latest advances in open-source projects and research are tracked and implemented. Best practices for LLM inference are extended and standardized within engineering workflows.
Requirements
Familiarity with running machine learning inference at large scale with high throughput and low latency is mandatory. Candidates must demonstrate this experience in real-world deployments.
Deep learning frameworks such as PyTorch must be understood alongside solid distributed systems and ML inference challenges. Mastery of these frameworks is essential for optimizing complex models.
U.S. citizenship, national origin, or visa status must comply with Anyscale's E-Verify participation and right-to-work requirements. This role requires adherence to export control and employment eligibility rules.
Practical notes
Role-based tasks involve fast iteration with product teams across batch and online workflows. employment eligibility and compliance with export and citizenship rules apply. Typical interview steps
Data interviews commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies give a take-home analysis. Expect questions about past projects and the business impact of your work. Interviewers often evaluate how you communicate uncertainty and business impact, not only the math. Bringing a clean write-up of a past analysis to the interview is well received.
Good to know
Distributed inference at scale relies on systems optimizations across frameworks and compilers. Ray powers scalable machine learning for companies such as OpenAI, Uber, Spotify, Instacart, and Cruise. ML infrastructure roles often work closely with open-source communities around LLM engines and deep learning frameworks. Compiler and runtime tools commonly include Triton, TVM, and MLIR for high-performance execution. GPU and CUDA experience is relevant for low-level optimization of inference workloads.
Questions to ask
Useful questions for the interview: what a typical week looks like, how work is assigned, what tools the team uses, and how feedback works. Asking how the role has changed recently and what the team wishes it had known when joining is also reasonable. Questions about the manager's priorities are especially valued.
Career growth
Data careers grow toward senior analyst, staff data scientist, or data engineering lead. Many professionals specialize in machine learning, analytics, or infrastructure. Cross-functional work with product and engineering teams becomes more important at senior levels. The field changes quickly, so continuous learning is part of the job. Professionals who can translate numbers into decisions tend to advance fastest.