MLOps Engineer
CognismPoland1w ago
Job description
About the role
Cognism is hiring an MLOps Engineer to join their growing team in Warsaw. This role focuses on building and maintaining the infrastructure that supports machine learning models in production environments. The engineer will work closely with data science and engineering teams to ensure reliable deployment, monitoring, and scaling of ML systems across the organization. Cognism provides business data solutions to companies worldwide, and this position plays a key part in improving how their data products are powered by machine learning and automation.
Key facts
What you'll do
- Design and maintain CI/CD pipelines for machine learning model deployment, testing, and automated retraining workflows
- Build and manage containerized environments for model serving using Docker and Kubernetes orchestration platforms at scale
- Monitor model performance in production and set up comprehensive alerting systems for data drift and model degradation
- Collaborate with data scientists to package, optimize, and validate trained models for scalable and efficient inference
- Develop infrastructure-as-code templates using Terraform or similar tools to provision and manage cloud resources for machine learning workloads
- Implement feature stores and reliable data pipelines that feed both training and real-time serving pipelines consistently
- Troubleshoot production issues related to model serving latency, throughput bottlenecks, and overall system availability under load
- Write automated tests for ML pipelines to ensure reproducibility and consistent reliability of model outputs over time
- Maintain version control and experiment tracking systems for managing model artifacts and configuration files effectively
- Optimize cloud resource usage to reduce operational costs while maintaining performance SLAs for machine learning services
- Document MLOps processes and runbooks to support team knowledge sharing and smooth onboarding of new team members
- Participate in on-call rotations to respond promptly to incidents affecting ML-powered products and services
- Coordinate with platform engineering teams to define standards and best practices for ML infrastructure development
- Contribute to the evolution of the company's MLOps strategy by identifying opportunities for process improvement
Requirements
- Experience with Python programming and writing production-grade code for building ML infrastructure, automation, and tooling
- Familiarity with container orchestration platforms such as Kubernetes or Docker Swarm for managing model serving at scale
- Understanding of machine learning fundamentals including training, evaluation, deployment lifecycles, model versioning, and reproducibility practices
- Hands-on experience with cloud platforms like AWS, GCP, or Azure for deploying, managing, and scaling ML workloads
- Knowledge of CI/CD tools and practices for automating ML model release, retraining, and deployment pipelines effectively
- Experience with version control systems, particularly Git, for managing infrastructure code and pipeline configuration files
- Familiarity with monitoring and observability tools for tracking model performance metrics and overall system health
- Strong problem-solving skills, attention to detail, and the ability to work independently on complex infrastructure challenges in production
Nice to have
- Experience with MLflow, Kubeflow, or similar MLOps platforms for experiment tracking, model registry, and pipeline orchestration
- Background in working with B2B data or data enrichment products and understanding their specific machine learning and analytics requirements
- Knowledge of SQL and data warehousing concepts for building reliable and scalable data pipelines
- Familiarity with infrastructure-as-code tools such as Terraform or Pulumi for cloud resource management and provisioning
Skills & tools
- Python for scripting, automation, orchestration, and building machine learning infrastructure components and microservices
- Docker and Kubernetes for containerizing applications, managing deployments, and orchestrating model serving workloads at scale
- CI/CD platforms such as Jenkins, GitLab CI, or GitHub Actions for automated ML pipeline execution
- Cloud providers including AWS, GCP, or Azure for deploying and scaling machine learning services
- Git for version control of infrastructure code, pipeline definitions, and configuration management files
- Monitoring tools such as Prometheus, Grafana, or Datadog for observability of machine learning systems
Practical notes
- This role is based in Warsaw and requires regular on-site presence at the Cognism office during standard working days
- The position is full-time with standard business hours, though some schedule flexibility may be available
- Candidates should be prepared to collaborate closely with cross-functional teams across data science and engineering departments on a daily basis
- Cognism offers a competitive benefits package, equity options, and a supportive environment for professional growth and development in MLOps