Senior Machine Learning Engineer
CloudflareHybrid4w ago
Job description
About the role
As a , you will play a role in the deployment and optimization of machine learning models across our extensive global network. Your work will span various applications, including open large language models (LLMs) and real-time voice processing. This position emphasizes deploying models on diverse GPU architectures and specialized accelerators while maintaining low latency and high reliability.
Key facts
What you'll do
- Develop and refine machine learning models tailored for our serverless inference platform, ensuring optimal performance and scalability.
- Establish comprehensive evaluation and benchmarking frameworks to monitor critical metrics such as throughput, latency, operational costs, and overall model performance.
- Implement advanced techniques to enhance inference speeds, including but not limited to quantization, caching, batching, and model compilation strategies.
- Collaborate with systems engineers to seamlessly integrate machine learning models into our distributed infrastructure, ensuring functionality.
- Oversee deployment workflows that encompass observability, regression testing, and safe rollout practices to mitigate risks.
- Translate customer requirements into scalable machine learning features specifically designed for Workers AI, enhancing user experience and satisfaction.
- Offer technical guidance and mentorship to team members, promoting adherence to production machine learning standards and best practices.
- Engage in cross-functional collaboration to align machine learning initiatives with broader organizational goals, fostering innovation and efficiency.
Requirements
- Proven experience in operating and optimizing machine learning models in a production environment, demonstrating a strong understanding of real-world applications.
- Proficiency in Python programming, along with expertise in machine learning frameworks such as JAX, TensorFlow, or PyTorch.
- Practical knowledge of inference optimization techniques, including runtime tuning and model compilation, to enhance model efficiency.
- Familiarity with serving runtimes or frameworks like vLLM, SGLang, TensorRT-LLM, llama.cpp, Triton, or ONNX Runtime, showcasing versatility in model deployment.
- Understanding of contemporary deep learning architectures, including vision models, large language models, multimodal models, speech models, and retrieval-augmented generation (RAG).
- Experience in optimizing models for specialized hardware accelerators or GPUs, ensuring high performance and reliability.
- Awareness of production-level concerns such as reliability, monitoring, and model evaluation, with a focus on maintaining operational excellence.
- Ability to connect the domains of systems engineering and machine learning, particularly in distributed systems or networking contexts.
- A demonstrated history of leading technical projects and providing mentorship to junior staff, fostering a culture of learning and growth.
Nice to have
- Contributions to open-source projects related to inference runtimes, model serving frameworks, or machine learning tools, showcasing a commitment to community and innovation.
- Familiarity with cloud-based machine learning services or platforms, enhancing deployment capabilities.
- Experience in working with real-time data processing systems, adding value to the role through practical insights.
Skills & tools
- Proficient in Python programming language.
- Experienced with machine learning frameworks such as PyTorch, TensorFlow, and JAX.
- Knowledgeable in serving runtimes and frameworks including SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, and llama.cpp.
- Familiar with various machine learning models, including LLMs, vision models, speech models, RAG, and embeddings.
- Competent in distributed systems, GPU optimization, and serverless infrastructure, ensuring effective model deployment.
Practical notes
- This position may require access to technology that is subject to U.S. export control regulations, and your ability to receive such technology without an export license may be a condition of your employment.
- The final stages of the interview process may necessitate an in-person visit to a Cloudflare office or hub, allowing for a deeper understanding of our culture and operations.
- Cloudflare is committed to being an equal opportunity employer and is dedicated to providing reasonable accommodations for applicants with disabilities. For assistance, please reach out to hr@cloudflare.com.
About the company
Cloudflare operates one of the largest networks in the world, providing security, performance, and reliability services to websites and internet applications. Founded by Matthew Prince, Lee Holloway, and Michelle Zatlyn in 2009, Cloudflare went public on the NYSE in September 2019.