Research Intern, Model Shaping
Job description
About the role
The Model Shaping team concentrates on adapting open foundation models for specific user needs and improving the efficiency of model training. You will engage directly with the research and engineering challenges of post-training methodologies, distributed training systems, and rigorous model evaluation. This internship provides a unique pathway to translate theoretical research into robust, production-ready features within a fast-paced cloud infrastructure environment. You will collaborate closely with engineers who are responsible for building the underlying systems that power large-scale AI training and inference. Your daily work will involve designing experiments that probe the capabilities and limitations of current models. The role emphasizes ownership, where your contributions will directly influence the direction of product-facing model improvements. You will be expected to communicate your progress clearly through both technical artifacts and written documentation.
Key facts
What you'll do
- Design and execute experiments to develop new methods in supervised learning, reinforcement learning, and preference optimization.
- Architect and build systems that enable efficient neural network training, focusing on both algorithmic innovation and distributed computing improvements.
- Construct frameworks and benchmarks for the reliable evaluation of foundation model performance across diverse tasks.
- Execute complex runs to verify research hypotheses and analyze large-scale training job outputs.
- Author scientific papers and technical blog posts that clearly communicate your findings to both technical and broader audiences.
- Translate experimental research outcomes into concrete features and improvements within Together AI products.
- Partner with distributed systems engineers to optimize training pipelines for speed, stability, and resource utilization.
- Implement monitoring and evaluation tools that provide actionable insights into model behavior during and after training.
- Investigate the trade-offs between model efficiency, accuracy, and alignment through systematic empirical study.
- Contribute to the creation of internal tools that streamline the workflow for model shaping and iteration.
Requirements
- You are currently enrolled in a Bachelor, Master, or Ph.D. program in Computer Science, Electrical Engineering, or a closely related discipline.
- You possess a solid grasp of Deep Learning and Machine Learning principles, including training dynamics and optimization techniques.
- You have practical experience with deep learning libraries such as JAX, PyTorch, or similar frameworks.
- You demonstrate a working knowledge of Transformer architectures and awareness of current trends in foundation models.
- You are proficient in writing clean, modular, and well-documented code that can be maintained by other engineers.
- You show the ability to debug complex systems and isolate root causes of performance regressions or failures.
- You communicate technical ideas effectively, both in writing and during collaborative discussions with team members.
- You are comfortable working in a fast-paced environment where priorities can shift based on research discoveries and product needs.
Nice to have
- Prior research background focused on efficient machine learning or foundation model training.
- Authorship of papers presented at top-tier venues such as ICLR, NeurIPS, ICML, EMNLP, or ACL.
- Familiarity with hardware acceleration strategies and advanced model optimization techniques.
- A demonstrated history of meaningful contributions to open-source machine learning software projects.
- Experience with large-scale training jobs on multi-GPU or TPU clusters.
- Knowledge of techniques for evaluating model alignment, safety, and robustness.
- Exposure to production deployment challenges for machine learning models.
Practical notes
- The internship compensation includes a housing stipend and additional benefits to support your living expenses.
- Final pay rates within the stated range are determined by individual skills, relevant experience, and geographic location.
- Together AI is committed to maintaining an inclusive environment and operates as an Equal Opportunity Employer.
- The engagement is structured as a full-time position running from September 14th to December 18th, spanning 12 to 16 weeks.
- This role is based in either San Francisco or Amsterdam, allowing for flexibility depending on candidate circumstances and team needs.
- Applicants should be available to start on the scheduled date to ensure seamless integration with ongoing research projects.
- Travel requirements are not expected as part of this role, given the remote-friendly nature of the arrangement.
- No visa sponsorship information is provided for international candidates at this time.