Member of Technical Staff
SpaceXAIUSA2w ago$180,000 - $440,000 USD
Job description
About the role
Become part of a dynamic team focused on achieving superhuman intelligence through the integration of various modalities including images, videos, audio, and text. In this position, you will engage in all phases of the development process, from gathering data and developing tokenizers to large-scale training, alignment, infrastructure setup, evaluation, and creating user-facing products. This role necessitates close collaboration with various teams, including those focused on pre-training, post-training, reasoning, data management, applied research, and product development, to enhance capabilities in multimodal reasoning, world modeling, tool utilization, and agentic systems.
Key facts
What you'll do
- Design and enhance large-scale distributed systems that facilitate multimodal pre-training, post-training, inference, data management, and tokenization at a petabyte scale.
- Construct high-throughput data pipelines for the acquisition, preprocessing, filtering, generation, decoding, loading, crawling, visualizing, and management of multimodal data including images, videos, audio, and text.
- Advance multimodal capabilities such as spatial-temporal compression, cross-modal alignment, world modeling, reasoning, emergent behaviors, and real-time video processing.
- Spearhead initiatives aimed at improving data quality, which includes human and synthetic curation, filtering techniques, analysis, and scalable pipelines tailored for trillion-parameter models.
- Create evaluation frameworks, internal benchmarks, reward models, and metrics that accurately reflect real-world usage, failure patterns, interactive dynamics, and human-AI collaboration.
- Innovate algorithms, modeling techniques, and approaches for hardware-software-algorithm co-design and scaling.
- Develop research tools, user-friendly interfaces, prototypes, demos, and comprehensive applications with rapid iteration cycles.
- Work across the entire stack from pre-training through supervised fine-tuning (SFT), reinforcement learning (RL), and post-training to facilitate reasoning, tool utilization, agentic behaviors, orchestration, and real-time interactions.
Requirements
- Proven experience in multimodal pre-training, post-training, or fine-tuning that involves vision, audio, video, or cross-modal applications.
- Advanced proficiency in Python, with substantial experience in at least one of the following: JAX, PyTorch, or XLA.
- A solid track record in building or optimizing large-scale distributed machine learning systems, including training and inference optimization, GPU utilization, and configurations involving multiple GPUs or TPUs, along with hardware co-design.
- Extensive experience in designing and managing data pipelines at scale, particularly for curation, filtering, generation, and quality analysis, especially with noisy real-world multimodal data.
- Strong foundation in evaluation design, benchmarks, reward modeling, or reinforcement learning techniques, particularly for interactive and agentic behaviors.
- A self-motivated approach that thrives in high-pressure environments.
- Readiness to take ownership of projects from start to finish, ensuring the delivery of comprehensive user experiences.
Nice to have
- A history of leading significant enhancements in model capabilities through improved data, modeling, algorithms, or scaling techniques.
- Familiarity with contemporary multimodal large language model (LLM) research, scaling laws, tokenizers, compression methods, reasoning, or agentic systems.
- Skills in Rust or C++ for components where performance is critical.
- Hands-on experience with large-scale orchestration tools such as Spark, Ray, or Kubernetes.
- Experience in developing full-stack tools, including efficient interfaces, real-time research demos and applications, or overseeing end-to-end product ownership.
- A keen interest in enhancing the user experience for interactive, real-time multimodal AI systems.
Skills & tools
- Proficient in Python, JAX, PyTorch, XLA
- Experience with distributed machine learning systems, including multi-GPU and TPU configurations
- Expertise in designing data pipelines at petabyte scale
- Knowledge of reward modeling and reinforcement learning techniques
- Familiarity with Rust and C++ (preferred)
- Experience with orchestration tools such as Spark, Ray, and Kubernetes (preferred)
Practical notes
- Compensation package includes equity in addition to the base salary.
- Comprehensive benefits package that covers medical, vision, and dental care.
- Access to a 401(k) retirement plan.
- Short-term and long-term disability insurance is provided.
- Life insurance and various discounts and perks are included.
- The organization has a flat structure, where taking initiative and delivering consistently can lead to leadership opportunities.
- Strong communication skills are essential for effective knowledge sharing with team members.