
Senior ML Research Scientist, Pegasus
Job description
About the role
The Senior ML Research Scientist, Pegasus role at TwelveLabs defines the research vision and execution strategy for Pegasus temporal segmentation and structured output generation across long-form video. You will operate at the intersection of foundational model research and production-scale video understanding, owning the technical roadmap for one of the core components of the Jockey system. This position requires deep expertise in multimodal learning, temporal modeling, and large-scale video analysis, with responsibility for translating complex customer needs into rigorous research objectives and measurable success criteria. You will work within a tightly integrated model organization where perception, reasoning, and agentic capabilities are developed cohesively, ensuring that advances in Pegasus directly enhance the capabilities of Jockey Core and downstream production workloads.
What you'll do
You will define and lead research initiatives for Pegasus focusing on temporal segmentation and structured output generation across long-form video. This includes designing scalable training strategies that span pre-training through reinforcement learning to improve reasoning and factual accuracy in video understanding. You will architect solutions for multi-hour context that maintain efficiency while preserving fine-grained temporal understanding, addressing one of the fundamental challenges in video intelligence. You will develop evaluation methodologies that correlate benchmark results with real-world customer workflows and production metrics, ensuring research progress translates into tangible product value.
You will partner with the Jockey Core team to integrate model advances into a cohesive agentic system with reliable segment-level reasoning, working at the boundary between model research and agent deployment. Collaboration with inference and serving teams will be essential to optimize model deployment, latency, and throughput under production constraints. You will lead data curation and quality efforts that ensure diverse, high-quality video corpora for robust model training, and define and own internal tools for experiment tracking, dataset management, and result visualization across research cycles. The role requires acting as a domain expert for product and engineering teams, translating customer requirements into research objectives, and communicating research progress through internal presentations, documentation, and external publications where appropriate. You will champion best practices in reproducibility, systematic evaluation, and scientific rigor across the model organization.
Requirements
You hold a PhD or equivalent experience in machine learning, computer vision, or a related quantitative field, with extensive experience in deep learning research focused on video or sequential multimodal models. You are proficient in designing and training large-scale neural networks, including experience with transformer-based architectures that form the foundation of modern video-language systems. A strong track record of publishing at top-tier conferences and journals in AI, computer vision, or multimedia demonstrates your ability to contribute original insights to the field. You are comfortable working in production environments and understand the trade-offs between research innovation and deployment constraints, which is critical for developing solutions that scale to millions of hours of video.
You have hands-on experience with training infrastructure, distributed computing, and large-scale data pipelines necessary for training advanced video understanding models. Excellent problem-solving skills enable you to decompose complex research problems into tractable sub-problems, while effective communication skills allow you to work with both technical and non-technical audiences in a fast-paced, ambiguous environment. The role demands fluency in Korean due to its Seoul, South Korea location, and candidates must be eligible to work under local visa regulations. Travel may be required for team meetings, partner discussions, or conferences, and deadlines for submission will be enforced per TwelveLabs personnel policies.
Nice to have
Experience with building and deploying production-grade video understanding systems will be valuable, along with familiarity with agentic workflows and reasoning across long-horizon tasks. Contributions to open-source video or multimodal model projects demonstrate engagement with the broader research community and practical experience with real-world video understanding challenges. Hands-on experience with reinforcement learning for language or vision-language models aligns with the responsibilities of designing training strategies that span pre-training through reinforcement learning to improve reasoning and factual accuracy in Pegasus.