Staff Machine Learning Ops Engineer
Job description
About the role
You will define and own the end-to-end technical vision and roadmap for Preply's machine learning platform, ensuring it meets the needs of growing ML and GenAI initiatives. In this role, you will set platform standards and design robust systems that enable ML teams to move from research to production quickly, safely, and cost-effectively. You will partner closely with Applied Science, Data, Product, and Engineering leadership to build a scalable, secure, and observable ML platform that powers multiple business lines. Your work will directly shape the infrastructure behind personalized learning, tutor and learner experiences, marketplace intelligence, content generation, automation, and future GenAI products. You will act as a technical leader and mentor, raising engineering standards and de-risking complex platform decisions across the organization. Your impact will be measured not only in the systems you build but in the speed, reliability, and confidence with which ML teams operate. You will play a critical role in making AI a scalable company-wide capability rather than a collection of isolated experiments.
Key facts
What you'll do
- Define the technical vision and roadmap for Preply's ML platform to support growing ML and GenAI adoption across teams, products, and business lines.
- Lead the architecture of platform capabilities across the full ML lifecycle, including experimentation, feature engineering, artifact management, training, deployment, monitoring, retraining, and governance.
- Design cloud-native infrastructure for distributed training and inference, including GPU-based environments, autoscaling, workload isolation, rollout strategies, and cost optimization.
- Set the technical direction for CI/CD for ML, embedding testing, validation, security, performance checks, and release confidence into deployment pipelines.
- Establish observability standards for ML systems, including model metrics, service health, alerts, drift detection, data quality, lineage, and business-impact monitoring.
- Lead the evolution of Preply's GenAI and LLM platform capabilities, including building LLM Gateway services, vector retrieval infrastructure, prompt experimentation, evaluation frameworks, latency-optimized inference, and reliable model-serving patterns.
- Partner with Applied Science and Data leads, Product leaders, and Engineering teams to align platform investments with experimentation velocity, cost efficiency, operational reliability, and user impact.
- Design platform abstractions, internal libraries, templates, and self-service tooling that help ML Scientists and engineers move faster without compromising reliability or security.
- Act as a technical multiplier across engineering by mentoring senior engineers, influencing architecture, raising standards, and guiding teams through complex platform decisions.
- Identify and eliminate bottlenecks in the path from ML research to production, making the platform easier, safer, and more efficient for all ML-powered product development.
Requirements
- 9+ years of engineering experience, with significant depth in large-scale ML, data, infrastructure, or platform systems.
- Proven ability to architect and scale production-grade ML platforms that support many teams, workflows, and ML use cases.
- Deep understanding of cloud-native architecture and end-to-end ML workflows, including experimentation, feature management, model versioning, training, deployment, monitoring, performance benchmarking, and optimization.
- Hands-on experience with modern ML infrastructure, including distributed training, GPU-based inference, autoscaling, and secure multi-tenant environments.
- Strong expertise in at least one major cloud provider, container orchestration, and infrastructure-as-code practices relevant to ML workloads.
- Experience designing and operating CI/CD pipelines tailored to ML, with a focus on testing, validation, security, and release confidence.
- Solid knowledge of observability for ML systems, including metrics, tracing, logging, alerts, drift detection, data quality checks, and lineage.
- Demonstrated ability to set technical standards, mentor engineers, and influence best practices across a large organization.
- Comfort working in a fast-paced, rapidly evolving environment where priorities can shift based on product and research needs.
Nice to have
- Experience with LLM Gateway, vector retrieval, and prompt management platforms.
- Background in building and operating evaluation frameworks for LLM outputs.
- Familiarity with latency-optimized inference and reliable model-serving patterns for GenAI workloads.