
Senior ML Research Scientist, Perception Models
Job description
Senior ML Research Scientist, Perception Models at Twelve Labs.
About the role
You will spearhead research initiatives focused on spatiotemporal modeling and multimodal representation learning within the Perception Models team at Twelve Labs. This role requires you to define complex research objectives that directly address the challenges of understanding video content through advanced computational methods. You will execute large-scale experiments designed to probe the fundamental limits of model understanding and representation. Your findings will be instrumental in shaping the direction of our video intelligence systems and influencing product roadmaps. You will deploy your discoveries and frameworks into production environments, ensuring that theoretical advances translate into tangible system improvements. This position involves close collaboration with engineering partners to integrate novel algorithms and capabilities into our core infrastructure. You will be responsible for creating training objectives, datasets, and evaluation benchmarks that set new standards in the field. Ultimately, your work will bridge the gap between cutting-edge research and real-world applications that derive meaning from visual and temporal data.
Key facts
What you'll do
- Investigate multimodal representations that bridge spatial and temporal data with semantic meaning to uncover deeper insights from video streams.
- Build models capable of processing information at various granularities, ranging from individual entities and regions to entire video clips, ensuring robustness across scales.
- Create training objectives, datasets, and evaluation frameworks for large-scale experimentation that push the boundaries of current technology.
- Incorporate signals from specialized vision tasks like segmentation, tracking, and detection into general-purpose model architectures to enhance perceptual accuracy.
- Collaborate with engineering and research departments to transition new capabilities into our production stack with efficiency and reliability.
- Design and implement end-to-end learning systems that leverage spatiotemporal modeling to solve complex real-world perception problems.
- Analyze massive video datasets to identify patterns, anomalies, and opportunities for novel model architectures and training strategies.
- Partner with cross-functional teams to align research goals with business objectives and user needs in dynamic market conditions.
- Explore innovative approaches to open-vocabulary understanding and region-level reasoning to expand the capabilities of our perception systems.
- Document research methodologies, experimental results, and theoretical insights to maintain a high standard of scientific rigor and reproducibility.
Requirements
- Demonstrated research history in multimodal learning, video understanding, or computer vision through publications or deployed systems.
- Deep technical knowledge in at least one of these areas: video foundation models, self-supervised or contrastive learning, embeddings, retrieval, temporal modeling, object-centric learning, detection, segmentation, or tracking.
- Proficiency in Python and deep learning frameworks such as PyTorch for implementing complex algorithms and models.
- Experience managing and executing rigorous, large-scale experiments that involve distributed computing and extensive hyperparameter tuning.
- A track record of impact through either academic publications in top-tier venues or deployed production systems that have improved performance metrics.
- Ability to navigate ambiguous technical challenges and work effectively across teams with diverse expertise and backgrounds.
- MS degree or equivalent industry experience in a production ML environment with a strong focus on real-world applications.
- Strong understanding of neural network architectures and their application to video and spatiotemporal data.
- Experience with data pipeline construction, model training, and deployment in production environments.
- Commitment to maintaining high standards of code quality, documentation, and experimental rigor.
Nice to have
- Experience with open-vocabulary vision or dense, region-level representations that enable models to understand a wide range of object categories.
- Background in large-scale video training, data curation, or infrastructure for model evaluation to support complex research initiatives.
- Experience moving research concepts into production environments and ensuring they meet strict performance and reliability standards.
- History of mentoring researchers or leading technical projects that demonstrate leadership and the ability to guide team members.
- Familiarity with modern deep learning frameworks and tools used for computer vision research and deployment.
- Prior work on systems that integrate multiple modalities, such as vision, language, and sensor data, to create comprehensive understanding.
Practical notes
- Benefits include an annual self-development budget of 1.4 million KRW, unlimited LLM token access for tech staff, and a 7.2 million KRW annual corporate card allowance for meals and transport.
- Hardware support includes a MacBook, a 700,000 KRW remote work equipment stipend, and a three-year replacement cycle for devices.
- Additional perks include health checkups for you and a family member, group insurance, flu shot coverage, and a two-week paid holiday break at the end of the year.
- We provide taxi fare support for late-night or weekend commutes and dinner coverage for office work past 7 PM.
- The position is based in Seoul, South Korea, with a hybrid work model that includes time in both Itaewon and Pangyo offices.
- Applicants must be eligible to work in South Korea under the current visa regulations.
- The interview process will involve technical assessments, coding challenges, and discussions with senior research staff.
- This role requires a commitment to collaborative work and active participation in the research community at Twelve Labs.
- The successful candidate will be expected to contribute to the broader research ecosystem through talks, publications, and internal knowledge sharing.
- The role demands a high level of self-motivation, discipline, and the ability to manage multiple research streams simultaneously.
- Candidates should be prepared to engage with cutting-edge technologies and contribute to the rapid iteration of our product offerings.
- The work environment is fast-paced and requires adaptability, resilience, and a passion for solving difficult problems in video perception.
- This position offers the opportunity to work with a talented team of researchers and engineers who are dedicated to advancing the state of the art in video understanding.
- The successful applicant will have the chance to influence the strategic direction of perception models and leave a lasting impact on the field.
- Travel is not required for this position, and all work can be conducted from our Seoul-based offices.
- The application deadline is strictly enforced, and late submissions will not be considered.
- Only candidates who fully meet the listed requirements will be contacted for the next stages of the hiring process.