Member of Technical Staff
Job description
About the role
The role centers on low precision reinforcement learning training and inference for a compact, mission focused team. Effective performance enables hands-on contribution within a flat, fast moving organization pursuing ambitious knowledge discovery. You will operate with a high degree of ownership over the entire lifecycle of reinforcement learning infrastructure, from foundational design to deployment and optimization. The work demands rigorous experimentation and analysis to ensure that systems perform reliably under demanding conditions. Success requires a blend of deep technical craftsmanship and the ability to translate abstract research goals into concrete, production ready solutions. You will collaborate closely with modeling experts to ensure the infrastructure accelerates scientific discovery rather than constraining it. Your contributions will directly impact the capability and efficiency of AI systems that operate in challenging environments.
Key facts
What you'll do
Reconfigure and optimize the inference stack to serve varied reinforcement learning scales, from small ablations to large production runs.
Design and tune the inference stack to minimize latency and maximize throughput across diverse reinforcement learning workloads.
Implement novel reinforcement learning techniques and algorithms in partnership with the modelling team to support real workloads and rapid experimentation.
Build, debug, and optimize large scale distributed systems to meet efficiency and reliability in demanding training and inference scenarios.
Handle large language model inference pipelines and deployment patterns with proven experience to ensure correctness under latency and throughput constraints.
Write Python, C++, and or Rust code while leveraging PyTorch, Jax, and CUDA to deliver robust infrastructure for model training and inference.
Demonstrate readiness to solve hard problems across the full software and hardware stack with a willingness to dive deep into complex issues.
Troubleshoot issues ranging from model code to low level kernel behavior to maintain resilient and scalable pipelines.
Leverage strong knowledge in quantization and numerics in LLM inference and training to improve accuracy and efficiency of systems.
Develop and optimize inference engines, such as SGLang and vLLM, to support high throughput serving and efficient resource utilization.
Requirements
The posting states a pay range of $180000 to $440000.
Hold a Bachelor's degree or equivalent experience as a strict eligibility requirement for the position.
Demonstrate the ability to build, debug, and optimize large scale distributed systems to ensure efficiency and reliability in demanding scenarios.
Possess proven experience in handling large language model inference pipelines and deployment patterns correctly.
Write Python, C++, and or Rust code competently while leveraging PyTorch, Jax, and CUDA to deliver infrastructure solutions.
Show readiness to solve hard problems across the full software and hardware stack without hesitation.
Demonstrate willingness to dive deep into complex issues from model code to low level kernel behavior for effective troubleshooting.
Apply strong knowledge in quantization and numerics in LLM inference and training to meet strict accuracy and efficiency standards.
Experience in developing inference engines like SGLang and vLLM is required to support high throughput serving.
Nice to have
Strong knowledge in quantization and numerics in LLM inference and training improves accuracy and efficiency.
Experience in developing inference engines, for example SGLang and vLLM, supports high throughput serving.
Practical notes
This role operates under equal opportunity employment practices. for specific eligibility and application steps.
Typical interview steps
Interviews in this field typically test judgment, communication, and fit as much as technical skill. Be ready to describe a challenge you faced, what you did, and what you learned. Most companies value honest answers over polished ones. Whatever the role, preparation is visible. Reviewing the company's product, the job description, and your own past work before the conversation is the strongest step.
Good to know
Work in this field centers on artificial intelligence systems that aim to understand complex environments. Engineers commonly use quantization and numerical methods to improve inference speed and stability. Development of inference engines such as SGLang and vLLM is common for high throughput serving. Strong communication skills help teams share knowledge accurately and concisely. Clear prioritization and a strong work ethic support success in fast moving research environments.
Questions to ask
Good questions to ask the employer in the interview: what does success look like in the first six months, how is the team structured, what is the current biggest challenge, and how are decisions made. Asking about growth paths and the review process is also well received. Employers expect questions, and good ones show preparation.
Career growth
Career growth comes from taking on harder problems and making your work visible. Look for opportunities to own outcomes, mentor others, and learn the business beyond your team. Growth in any career comes from scope, results, and reputation. Take on work that is slightly uncomfortable, and make sure others can see the outcomes.