Member of Technical Staff - Research & Post-training
Job description
Member of Technical Staff - Research & Post-training at Preference Model.
About the role
Preference Model is at the forefront of automated machine learning research engineering. We are tackling the challenges posed by existing frontier models, which often struggle when applied to real-world machine learning tasks. Our primary focus is on creating high-quality reinforcement learning (RL) training environments that accurately reflect the complexities of real-world scenarios. As part of our team, you will contribute to pioneering research in self-directed learning, specifically in the realm of post-training for large language models. This position merges research and engineering, allowing you to implement innovative strategies and influence the direction of our research initiatives.
Key facts
What you'll do
- Conduct training and evaluation of models within our proprietary RL environments to ensure data quality, identify gaps in task coverage, and enhance the feedback loop between environment design and model performance.
- Design and optimize our RL training infrastructure, focusing on training abstractions and distributed experiment management using frameworks such as Verl, OpenRLHF, or similar technologies.
- Create, implement, and assess training environments, evaluation metrics, and methodologies tailored for RL agents.
- Profile and enhance training processes from start to finish, including data loading and reward computation, to maximize throughput and reduce the research iteration cycle.
- Collaborate closely with cross-functional teams to align research objectives with engineering capabilities.
- Stay informed about the latest advancements in post-training research and translate theoretical concepts into practical implementations.
- Develop and maintain documentation for training environments and methodologies to ensure reproducibility and facilitate knowledge sharing.
- Engage in regular code reviews and contribute to best practices for code structure and organization in RL training projects.
Requirements
- Proven experience in managing end-to-end post-training pipelines for large language models, specifically those with a minimum size of 7 billion parameters.
- Strong proficiency in programming languages such as Python, along with experience using machine learning libraries like PyTorch or JAX.
- Familiarity with at least one contemporary reinforcement learning training framework.
- Experience in building and maintaining machine learning infrastructure capable of scaling to meet research demands.
- Ability to evaluate model outputs and develop reward or evaluation signals effectively.
- Excellent communication skills, with the ability to articulate complex concepts clearly and collaborate with diverse teams.
Nice to have
- A background in evaluating and refining model outputs, particularly in developing reward mechanisms.
- A proactive approach to staying updated on the latest post-training research and the ability to convert academic papers into functional code.
- Strong opinions (that can be adjusted) regarding the organization of RL training code for enhanced reproducibility and rapid iteration.
- A balanced mindset that merges research exploration with engineering discipline.
- Exceptional systems design capabilities and the ability to communicate effectively with both technical and non-technical stakeholders.
Skills & tools
- Python
- PyTorch or JAX
- Reinforcement Learning frameworks (e.g., Verl, OpenRLHF)
- Machine Learning infrastructure management
Practical notes
Preference Model is committed to fostering a diverse and inclusive work environment. We encourage individuals from all backgrounds to apply, even if they do not meet every single requirement listed. We believe that adaptability, along with strong communication and collaboration skills, are key factors in the success of our research initiatives.
Our team values the contributions of each member and offers a competitive compensation package that includes both cash and equity, placing us above the 90th percentile in the industry. You will enjoy a high degree of ownership and autonomy in a dynamic startup atmosphere, working alongside some of the brightest minds in machine learning. Additional benefits include comprehensive health, vision, and dental coverage, a 401K match, daily onsite lunches, and weekly snack deliveries. We also provide visa sponsorship and relocation support to help you transition smoothly to our San Francisco office.
Join us at Preference Model, where your expertise can help shape the future of machine learning and contribute to advancements in AI technology.