Sr Machine Learning Engineer, AI Research
Job description
About the role
You will design, train, and evaluate machine learning models across a range of research and applied AI initiatives that directly advance the Cribl AI Platform for Telemetry. You will run rapid, iterative experiments to test hypotheses and surface insights that drive measurable model improvements and inform product decisions. Collaborate closely with researchers and engineers to translate cutting-edge academic advances into practical, production-ready systems that solve real customer problems in telemetry and observability. You will build and maintain robust ML pipelines for data ingestion, feature engineering, model training, and evaluation with reliability and scalability as core principles. Optimize model performance through fine-tuning, hyperparameter search, and architecture experimentation to meet defined accuracy, latency, and operational targets. Contribute to a culture of rigorous experimentation by tracking results, documenting findings, and sharing learnings with the broader team to accelerate collective progress. Stay current with the latest developments in ML and AI research, and proactively identify opportunities to apply novel techniques to Cribl's telemetry and security observability use cases. This position may require stand-by, on-call, or off-hours duties during critical research or deployment milestones to ensure continuity and rapid response.
Key facts
What you'll do
Design, train, and evaluate machine learning models across a range of research and applied AI initiatives that directly impact the Cribl AI Platform for Telemetry.
Run rapid, iterative experiments to test hypotheses and surface insights that drive model improvements and inform product decisions.
Collaborate closely with researchers and engineers to translate cutting-edge academic advances into practical, production-ready systems.
Build and maintain robust ML pipelines for data ingestion, feature engineering, model training, and evaluation with reliability and scalability as core principles.
Optimize model performance through fine-tuning, hyperparameter search, and architecture experimentation to meet defined accuracy, latency, and operational targets.
Contribute to a culture of rigorous experimentation by tracking results, documenting findings, and sharing learnings with the broader team to accelerate collective progress.
Stay current with the latest developments in ML and AI research, and proactively identify opportunities to apply novel techniques to Cribl's telemetry and security observability use cases.
This position may require stand-by, on-call, or off-hours duties during critical research or deployment milestones to ensure continuity and rapid response.
Work closely with development partners and key stakeholders to iteratively design, develop, and deliver products and surfaces that will delight our customers using AI capabilities.
Leverage strong analytical and problem-solving skills to break down complex telemetry challenges into well-scoped ML problems.
Partner with infrastructure and platform teams to ensure ML solutions integrate smoothly into existing data and tooling landscapes.
Explore and prototype emerging AI/ML methods that can provide a differentiated customer experience in observability and security operations.
Document models, experiments, and decisions to ensure reproducibility, transparency, and alignment with best practices.
Participate in code reviews, design discussions, and cross-functional planning to elevate the overall quality and maintainability of ML deliverables.
Requirements
Bachelor's degree in Computer Science, Mathematics, Statistics, or a related field with 4+ years of industry or research experience (Master's or PhD a plus).
Deep hands-on experience training and evaluating ML models, including language models, in real-world settings.
Strong proficiency in Python and ML frameworks such as PyTorch or TensorFlow for building and deploying models.
Familiarity with MLOps tooling and infrastructure (e.g., MLflow, Weights & Biases, Kubeflow, or similar) to enable reliable experimentation and deployment.
Solid understanding of modern NLP, computer vision, and/or reinforcement learning techniques relevant to telemetry and observability problems.
Strong ability to move fast without sacrificing rigor; you know when to prototype and when to productionize based on empirical results.
Excellent communication skills with the ability to clearly present experimental results to both technical and non-technical stakeholders.
Ability to work effectively in a remote-first, collaborative environment with distributed teams across multiple time zones.
Willingness to adhere to Cribl's principles of curiosity, collaboration, and continuous learning in service of customers and the product.
Comfort operating in an environment where responsibilities may shift quickly and priorities are clarified through experimentation and data.
Committed to maintaining high standards for model quality, security, privacy, and operational reliability in all delivered work.