Senior Machine Learning Engineer, Jockey Core
Job description
About the role
This role defines the serving strategy and infrastructure for Jockey Core at TwelveLabs, from benchmark design through Blackwell-scale deployment. You will own the critical path that turns model architecture into real-time agent responses across millions of video queries. You will build load tests that mirror production traffic patterns to isolate latency, throughput, and cost tradeoffs. You will apply advanced inference optimization techniques to maximize utilization of the most advanced accelerators in the world. You will work closely with model researchers to align serving behavior with training objectives and evaluation metrics. You will own end-to-end reliability, monitoring, and performance for the component that reasons over all retrieved video and image context. You will translate product requirements into low-level serving configurations that enable new agent capabilities without re-integration.
Key facts
What you'll do
Design and maintain benchmarks and load tests that replay short prompts, frequent round-trips, and long tool-output contexts for Jockey Core.
Measure TTFT and inter-token latency independently to isolate pipeline and model-level performance issues.
Apply inference optimization methods such as quantization, batching, scheduling, and disaggregated prefill and decode.
Implement speculative decoding and other acceleration techniques to reduce end-to-end agent latency.
Scale serving infrastructure to production levels on Blackwell and other cutting-edge GPU architectures.
Collaborate with model training teams to ensure serving behavior reflects training objectives and evaluation signals.
Define and track reliability, monitoring, and alerting metrics for the core reasoning model in production.
Translate product requirements into concrete serving configurations that unlock new agentic workflows.
Own capacity planning and cost optimization for high-volume video and image understanding workloads.
Drive experiments that compare different engine choices and infrastructure layouts for optimal quality and efficiency.
Partner with product and research to prototype new agent features and validate their feasibility on existing infrastructure.
Document serving patterns, failure modes, and optimization strategies for cross-team consumption and reuse.
Requirements
You hold a Bachelor's, Master's, or PhD in Computer Science, Machine Learning, or a related technical field.
You have 5+ years of production machine learning or deep learning engineering experience, with at least 3 years focused on serving or inference optimization.
You are proficient in Python and in building performance-critical systems.
You have hands-on experience with modern GPU architectures such as NVIDIA Blackwell or comparable accelerators.
You understand large language model inference patterns including prefill, decode, batching, and speculative decoding.
You have used or built benchmarking and load testing tools for LLM serving workloads.
You are comfortable working in a fast-paced, research-to-production environment with rapidly evolving models.
Nice to have
Experience with agentic systems and retrieval-augmented workflows is preferred.
Familiarity with multimodal models and video understanding pipelines is a plus.
Knowledge of training-to-serving alignment and evaluation frameworks is beneficial.
About the team
The Cognition Models team owns the models that turn video into structured understanding and reasoning: Pegasus, our video-language model, and Jockey Core, the reasoning LLM behind Jockey. In the model stack we sit between Perception Models (embeddings and retrieval) and the agent system - taking what's retrieved and producing structured understanding and the reasoning to act on it. We focus on multimodal systems with high instruction-following capability and complex, hierarchically structured outputs. Our work spans training infrastructure from pre-training to RL, temporal segmentation and structured metadata extraction, large-scale inference and serving systems, data-curation and evaluation pipelines, and building Jockey Core. We ship products with real-world value rather than doing research in isolation, working as a goal-oriented, cross-functional team of ML researchers and engineers - using the most advanced compute in the world, including NVIDIA B300s, to accelerate the research-to-production cycle.
About Jockey Core
Jockey Core is the reasoning LLM at the center of Jockey - the model that decomposes a query, decides what to retrieve and segment, and reasons over the results into an answer you can act on. It sits in the critical path of every agent step, so its quality, latency, and cost directly shape what Jockey can do. Jockey Core is a model we own and serve end to end, and we improve it continuously so Jockey's quality compounds with every release.
About TwelveLabs
Video is 90% of the world's data. Most of it is invisible to machines. TwelveLabs builds the intelligence layer to change that. Our multimodal AI models understand video the way humans do - across sight, sound, and motion - and power production-scale AI workloads across media, entertainment, sports, security, and government. We have raised more than $210 million from NEA, Radical Ventures, Amazon, NVIDIA, Snowflake, Databricks, Index Ventures, NAVER Ventures, Korea Investment Partners, Quadrille Capital, Red Bull Ventures, and AI pioneers including Fei-Fei Li, Silvio Savarese, and Alexandr Wang.
We are a global company, headquartered in San Francisco with offices in Seoul, New York, and London, and employees around the world. We believe the differences in our cultural, educational, and life experiences make our products stronger. Building technology that understands the world in all its complexity requires people who see it from every angle. We are looking for individuals who are driven by hard problems and want their work to matter. Come build it with us.