Senior Machine Learning Engineer, DevOps/SRE
Job description
About the role
Roku is seeking a seasoned Senior Machine Learning Engineer with a strong background in DevOps and Site Reliability Engineering (SRE) to join our innovative team in San Jose, California. This position is in enhancing our machine learning capabilities while ensuring the reliability and scalability of our systems. As a key contributor, you will collaborate with cross-functional teams to develop and deploy machine learning models that power our streaming services, ultimately enriching the user experience. Your expertise will help shape the future of our platform as we continue to lead the industry in streaming technology.
Key facts
What you'll do
- Design and implement machine learning models that enhance the performance of Roku's streaming services.
- Collaborate with data scientists and software engineers to integrate machine learning algorithms into production systems.
- Develop and maintain CI/CD pipelines to automate the deployment of machine learning models and related infrastructure.
- Monitor and optimize the performance of machine learning systems, ensuring high availability and reliability.
- Conduct experiments to evaluate the effectiveness of various machine learning approaches and refine models based on findings.
- Work closely with product managers to understand user requirements and translate them into technical specifications for machine learning solutions.
- Implement best practices for model versioning, testing, and monitoring to ensure continuous improvement of deployed models.
- Participate in code reviews and provide mentorship to junior engineers, fostering a culture of learning and collaboration within the team.
- Collaborate with the SRE team to ensure that machine learning services are resilient and scalable, addressing any operational issues as they arise.
- Stay updated with the latest advancements in machine learning and SRE practices, applying new knowledge to enhance existing systems.
- Contribute to the documentation of processes, architectures, and models to facilitate knowledge sharing within the organization.
- Engage in troubleshooting and debugging of machine learning systems, identifying root causes and implementing effective solutions.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
- A minimum of 5 years of experience in machine learning engineering, with a focus on deploying models in production environments.
- Strong proficiency in programming languages such as Python, Java, or Scala, with experience in machine learning libraries like TensorFlow or PyTorch.
- Solid understanding of DevOps principles and practices, including CI/CD, containerization (Docker), and orchestration (Kubernetes).
- Experience with cloud platforms such as AWS, Google Cloud, or Azure, particularly in deploying machine learning applications.
- Familiarity with data engineering concepts and tools for data ingestion, transformation, and storage.
- Excellent problem-solving skills and the ability to work independently as well as part of a team.
- Strong communication skills, capable of conveying complex technical concepts to non-technical stakeholders.
- Proven track record of delivering high-quality software solutions on time and within scope.
Nice to have
- Experience with big data technologies such as Apache Spark or Hadoop.
- Knowledge of MLOps practices for managing the lifecycle of machine learning models.
- Familiarity with monitoring and observability tools like Prometheus or Grafana.
- Contributions to open-source machine learning projects or active participation in relevant communities.
- Previous experience in the streaming media industry or related fields.
Skills & tools
- Programming Languages: Python, Java, Scala
- Machine Learning Frameworks: TensorFlow, PyTorch
- DevOps Tools: Docker, Kubernetes, Jenkins
- Cloud Platforms: AWS, Google Cloud, Azure
- Data Engineering: Apache Spark, Hadoop
- Monitoring Tools: Prometheus, Grafana
Practical notes
- This position is based in San Jose, California, and offers a competitive salary based on experience.
- Candidates requiring visa sponsorship should apply, as Roku is open to supporting qualified individuals.
- The role may involve occasional on-call responsibilities to ensure system reliability.
- Interested applicants can find more details and apply through the Roku careers page.
Join Roku and be part of a team that is revolutionizing the way people experience entertainment. Your expertise in machine learning and DevOps will play a crucial role in shaping the future of streaming technology.