Senior ML Research Scientist, Jockey Core
Job description
Senior ML Research Scientist, Jockey Core at Twelve Labs.
About the role
You will lead model-efficiency and post-training research for Jockey Core, owning the research roadmap that makes a high-quality reasoning model efficient enough to serve in production without losing what makes it good. You will drive model compression initiatives such as structured pruning, quantization (PTQ/QAT), distillation, and recovery fine-tuning to build reasoning models that keep their quality at a production-efficient size. You will design rigorous evaluation protocols and data curation strategies for post-training, ensuring that efficiency gains are measured and aligned with real-world agent behaviors. You will collaborate closely with the Pegasus and Jockey Core teams to translate retrieval and perception signals into structured training signals for compression and fine-tuning. You will partner with production and infrastructure engineers to prototype and deploy efficient model variants on real hardware with strict latency, memory, and cost targets. You will analyze failure cases from live agent traces to identify whether inefficiencies stem from representation, training, or architectural choices. You will own experiments that balance accuracy, throughput, and cost, and clearly communicate trade-offs to both technical and non-technical stakeholders. You will contribute to internal tools and evaluation dashboards that make model behavior and efficiency improvements observable across the Jockey stack.
Key facts
What you'll do
Drive model compression and efficiency research - structured pruning, quantization (PTQ/QAT), distillation, and recovery fine-tuning - building reasoning models that keep their quality at a production-efficient size.
Design rigorous evaluation protocols and data curation strategies for post-training, ensuring that efficiency gains are measured and aligned with real-world agent behaviors.
Collaborate closely with the Pegasus and Jockey Core teams to translate retrieval and perception signals into structured training signals for compression and fine-tuning.
Partner with production and infrastructure engineers to prototype and deploy efficient model variants on real hardware with strict latency, memory, and cost targets.
Analyze failure cases from live agent traces to identify whether inefficiencies stem from representation, training, or architectural choices.
Own experiments that balance accuracy, throughput, and cost, and clearly communicate trade-offs to both technical and non-technical stakeholders.
Contribute to internal tools and evaluation dashboards that make model behavior and efficiency improvements observable across the Jockey stack.
Champion best practices in reproducibility, baselines, and experimental hygiene across the model and agent teams.
Support the definition of long-term model architecture strategies for Jockey Core within the constraints of production environments.
Mentor junior researchers and engineers on compression techniques, evaluation design, and responsible AI practices.
Participate in cross-functional discussions that shape product requirements based on model capabilities and limitations.
Contribute to publications, internal technical documentation, and open-source components where appropriate.
Engage with the broader research community on model efficiency, post-training, and multimodal reasoning topics.
Translate high-level product goals into concrete research experiments with measurable success criteria.
Ensure all research activities are aligned with the company's production constraints and long-term model roadmap.
Requirements
8+ years of experience in machine learning research or a related field with a strong track record of delivering production-grade models.
Strong expertise in deep learning for NLP or multimodal models, including training, fine-tuning, and optimization techniques.
Demonstrated experience with model compression techniques such as pruning, quantization, and distillation in production settings.
Proficiency in PyTorch and modern training and inference frameworks, with hands-on experience in deploying models at scale.
Solid understanding of large language models and reasoning-oriented architectures, including attention mechanisms and emergent capabilities.
Experience with building and maintaining evaluation frameworks and data pipelines for model performance and safety.
Strong grasp of software engineering best practices, including version control, testing, and modular code design.
Ability to work cross-functionally with product, infrastructure, and research teams to align on goals and constraints.
Nice to have
Experience with video or multimodal models and tokenization strategies.
Familiarity with serving infrastructure and hardware-aware optimization.
Contributions to open-source model compression or post-training tools.
Practical notes
in the provided source.