Senior Machine Learning Operations Engineer
MercuryUSA1w ago
Machine LearningOperationsEngineeringremotecurated-jd
Job description
Senior Machine Learning Operations Engineer at Mercury
About the role
Mercury is expanding its use of machine learning for risk assessment, with models increasingly influencing critical real-time decisions. The Machine Learning Platform team is tasked with creating a streamlined process for deploying and monitoring these models in production, enabling faster development cycles and detailed performance insights.
Key facts
What you'll do
- Develop and maintain the infrastructure for real-time model inference, ensuring low latency and high availability for the risk decision engine.
- Manage the systems for model deployment, including version control, automated checks for performance and bias, and phased rollout strategies.
- Implement comprehensive monitoring for deployed models, covering availability, response times, error rates, and drift detection to signal retraining needs.
- Collaborate closely with Risk Data Science to ensure a smooth transition of models from development to production operation under the MLP team's purview.
- Build capabilities for A/B testing models, such as champion/challenger setups and canary deployments, and integrate model explainability features.
Requirements
- Possess at least 5 years of experience in machine learning engineering, backend software development, MLOps, or a closely related discipline.
- Demonstrate experience deploying, serving, and operating machine learning models in production environments that require low latency and high availability.
- Exhibit strong backend engineering skills in Python, including experience with API frameworks like FastAPI or Flask.
- Have a solid understanding of model deployment and lifecycle management tools, including model registries, CI/CD pipelines for models, versioning, and staged rollout patterns.
- Proven ability to build observability and alerting for production services, covering metrics like latency, errors, and model-specific signals such as drift.
- Be comfortable working with data infrastructure that supports machine learning, including SQL, low-latency data stores (e.g., Redis, DynamoDB), and streaming platforms (e.g., Kafka, Kinesis).
Nice to have
- Familiarity with modern data stack components like Snowflake, dbt, Dagster, or Airflow.
- Experience working within regulated, audit-sensitive, or compliance-focused environments.
- Exposure to functional programming languages or an interest in working with a stack that includes Haskell, React, and TypeScript.
Skills & tools
- Python
- FastAPI or Flask
- Model Registries
- CI/CD for Models
- Observability and Alerting Tools
- SQL
- Redis, DynamoDB, or equivalent
- Kafka, Kinesis, Redpanda, or equivalent
Practical notes
- US employees (any location) base salary range: $166,600 - $208,300.
- Canadian employees (any location) base salary range: CAD 157,400 - 196,800.
- Total compensation includes base salary, equity, and benefits. Salary and equity ranges are competitive and based on industry data, candidate experience, expertise, location, and internal equity.
- Mercury is an Equal Employment Opportunity employer committed to diversity and inclusion. Reasonable accommodations are available for applicants with disabilities.