Staff AI Engineer, Model Post-Training and Alignment
Job description
About the role
The role centers on architecting and executing large model post-training and alignment strategies that directly enhance capability and safety. You will own the design of end-to-end training loops, from data strategy and reward modeling to reinforcement learning optimization. This position requires deep collaboration with research and engineering teams to translate alignment objectives into scalable experiments. You will be responsible for ensuring that post-training processes improve controllability, reasoning, and domain adaptation in production environments. The position involves close partnership with data scientists and infrastructure engineers to standardize training frameworks. You will also evaluate emerging alignment techniques and assess their applicability to real-world deployment constraints. Ultimately, your work will shape how models behave and perform for millions of users on the OKX platform.
Key facts
What you'll do
- Lead the conception and execution of large language model post-training pipelines, supervising fine-tuning, preference optimization, and reinforcement learning workflows.
- Implement advanced alignment methodologies including Direct Preference Optimization and Generalized Reward Policy Optimization with rigorous experimentation.
- Curate and structure high-quality domain-specific datasets, defining curation rules and augmentation techniques to maximize task performance.
- Conduct post-training for specialized small models, covering architecture selection, dataset construction, and optimization strategy formulation.
- Develop and maintain Reward Models that provide reliable feedback signals for alignment and downstream optimization tasks.
- Establish closed-loop RLAIF systems that integrate AI feedback into training cycles to automate and refine model behavior.
- Drive inference optimization initiatives, leveraging serving frameworks such as vLLM and SGLang to reduce latency and increase throughput.
- Partner with infrastructure teams to ensure training and inference pipelines are robust, scalable, and aligned with operational standards.
- Analyze experiment results, diagnose failure modes, and iterate on training configurations to continuously improve model quality.
- Document methodologies, results, and best practices to enable knowledge sharing and reproducibility across teams.
- Collaborate with product teams to align post-training objectives with business goals and user safety requirements.
- Mentor junior engineers by providing code reviews, technical guidance, and constructive feedback on model development.
- Track and report on key performance indicators related to model accuracy, safety, and efficiency throughout the training lifecycle.
- Contribute to open source tools and internal frameworks that accelerate post-training research and deployment.
- Ensure all work adheres to company policies, ethical guidelines, and regulatory expectations for AI systems.
- Act as a technical owner for post-training initiatives, translating ambiguous problems into concrete execution plans.
Requirements
- Hold a Bachelor's degree, Master's degree, or PhD in Computer Science, Machine Learning, or a closely related technical field.
- Demonstrate substantial experience with large language model post-training, including supervised fine-tuning, preference learning, and reinforcement learning.
- Show proven ability in implementing and tuning Direct Preference Optimization and Generalized Reward Policy Optimization methods.
- Exhibit strong skills in building and refining Reward Models for alignment and optimization workflows.
- Have hands-on experience with Reinforcement Learning from AI Feedback in closed-loop training systems.
- Proficiency in optimizing inference using high-performance serving frameworks such as vLLM and SGLang is required.
- Solid understanding of data curation, dataset construction, and augmentation strategies for domain adaptation is essential.
- Experience working across the full machine learning lifecycle from experiment design to production deployment is mandatory.
Practical notes
- Location is specified as APAC.
- Engagement type is to be confirmed from the source.
- No compensation details are provided in the source.
- No hours, travel, visa, or deadline information is available beyond location and engagement notes.
The role is positioned at the intersection of research and engineering, focusing on taking large models from experimental training stages to reliable, high-performance production systems. Candidates should be comfortable working in a fast-paced environment where technical rigor and clear communication are essential. You will be expected to contribute not only to immediate project goals but also to the long-term evolution of the team's capabilities and practices. The position demands ownership, curiosity, and a methodical approach to problem-solving. Success will be measured by improvements in model quality, alignment, and efficiency that can be observed in real-world usage. Collaboration across teams will be a constant, requiring adaptability and a willingness to share knowledge. The work involves both deep technical challenges and practical constraints that must be balanced to deliver effective solutions. This role is ideal for engineers who want to shape the future of AI-driven financial services in the crypto industry.