Senior Machine Learning Operations Engineer
ParamountUSA1w ago
Machine LearningOperationsEngineeringremotecurated-jd
Job description
Senior Machine Learning Operations Engineer at Paramount
About the role
Paramount is seeking a Senior Machine Learning Operations Engineer to oversee the operational aspects of our personalization and recommendation machine learning systems. In this role, you will ensure the reliability of our models, which are retrained and deployed daily through automated processes. Your primary focus will be on monitoring system performance and swiftly addressing any issues that arise.
Key facts
What you'll do
- Manage model traceability to ensure every production model has clear documentation regarding its training data, code, validation processes, and performance metrics. You will assess and suggest tools for version control, metadata management, and model registries, collaborating with machine learning engineers to promote their use.
- Establish comprehensive monitoring systems to track data flow, feature distribution stability, model performance metrics, and service latency in accordance with service level agreements. You will take ownership of this process, ensuring that issues are identified without relying solely on upstream teams.
- Collaborate with data engineering teams to address data quality challenges, identify drift in upstream data sources, and maintain the reliability of features.
- Proactively monitor for issues, tracking drift over time to identify and address performance degradation before it becomes critical, and ensuring feature freshness to prevent cascading problems.
- Develop diagnostic tools to quickly identify root causes of issues when recommendations appear inaccurate, ensuring that relevant context is logged at every stage and creating dashboards to visualize this information.
- Lead incident response efforts for machine learning systems, maintaining rollback procedures and predefined strategies for hotfixes, while also managing automated checks to prevent faulty deployments. Conduct post-mortems to identify and rectify gaps in processes.
- Work with machine learning engineers, data engineers, and stakeholders to determine which metrics to collect post-deployment and understand their significance.
Requirements
- At least 5 years of experience in machine learning engineering, applied machine learning, or a similar role, with a proven track record in monitoring, reliability, deployment, or incident management.
- Experience in building or managing model registries, machine learning monitoring systems, or production machine learning pipelines.
- Comprehensive understanding of machine learning systems from end to end, including the implications of stale features or shifted distributions.
- Strong SQL skills with the ability to analyze data distributions, feature health, and model behavior.
- Ability to collaborate effectively with DevOps and Platform teams to define infrastructure requirements without needing to manage the infrastructure directly.
Nice to have
- Experience in operating recommendation or personalization systems at scale.
Skills & tools
- Machine Learning Operations (MLOps)
- Monitoring and diagnostics tools
- SQL
- Collaboration with DevOps and Platform Engineering
Practical notes
- The salary range for this position is between $139,200.00 and $208,800.00, depending on various factors such as location, experience, and education.
- Benefits include medical, dental, vision, 401(k), life insurance, disability coverage, tuition assistance, and paid time off. This position is also eligible for bonuses.
- Paramount is an equal opportunity employer and is committed to creating an inclusive environment for all employees. If you require accommodations due to a disability, please reach out via the provided contact methods.