Senior Applied Scientist, Large Language Models
Job description
About the role
Patsnap is actively seeking a Senior Applied Scientist focused on Large Language Models to join its dynamic team in Shanghai. In this capacity, you will own the design and delivery of advanced language model features that directly power knowledge-intensive products and strategic initiatives. Your primary responsibility will be to act as the crucial bridge between cutting edge research and robust, scalable production systems. You will manage the complete lifecycle of model development, rigorously overseeing experimentation, validation, and controlled deployment. A core part of this role involves enhancing reasoning capabilities, long context understanding, and precise data extraction to meet evolving business demands. You will investigate model failures, diagnose root causes, and implement targeted corrective actions to ensure high quality outputs. Furthermore, you will collaborate closely with domain experts to align technical capabilities with specific commercial objectives. Ultimately, your work will define the technical standards and help shape the internal AI roadmap for the organization.
Key facts
What you'll do
- Architect and iterate on language model capabilities tailored for specific business workflows and enterprise requirements.
- Drive improvements in model reasoning logic, long-context window handling, and accurate structured data extraction pipelines.
- Design and implement supervised fine-tuning regimens, preference optimization procedures, and knowledge distillation projects.
- Build and maintain comprehensive evaluation frameworks that track accuracy, factuality, safety compliance, latency profiles, and cost metrics.
- Develop end to end data preparation pipelines, experiment tracking systems, and standardized model assessment protocols.
- Conduct in depth error analysis of model outputs to identify systemic issues and formulate effective remediation plans.
- Critically assess emerging research papers and determine the practical applicability of novel techniques within production constraints.
- Partner closely with software engineers to optimize model serving infrastructure and ensure reliable large scale deployment.
- Translate ambiguous product concepts into concrete technical specifications in collaboration with product managers and stakeholders.
- Contribute to the definition of internal best practices, tooling standards, and the strategic direction of AI initiatives.
- Mentor junior engineers and researchers, fostering a culture of rigorous experimentation and continuous learning.
- Evaluate and integrate new model architectures, tool using systems, and agent frameworks where they provide measurable advantages.
- Establish robust data flywheels and feedback mechanisms to facilitate ongoing model refinement and performance gains.
- Ensure all solutions adhere to strict latency, cost, and operational benchmarks required for enterprise grade products.
Requirements
- Hold a Master's degree or PhD in Computer Science, Artificial Intelligence, Machine Learning, Natural Language Processing, or a closely related quantitative field.
- Demonstrate a strong foundational background in modern machine learning techniques and theoretical principles.
- Possess hands on experience in building, fine tuning, or extensively modifying large language models in real world scenarios.
- Show deep familiarity with Transformer based architectures, training methodologies, fine tuning strategies, and inference mechanisms.
- Exhibit proficiency in at least two specialized areas such as LLM post training and alignment, systematic benchmarking, retrieval augmented generation, long context modeling, advanced information extraction, complex logical reasoning and planning, or model compression and inference optimization.
- Write clean, efficient Python code and have substantial practical experience using the PyTorch deep learning framework.
- Proven ability to own and deliver end to end algorithmic projects, from initial problem scoping through to production grade implementation.
- Communicate effectively with cross functional partners, articulating technical concepts to both technical and non technical audiences.
- Understand the constraints and trade offs involved in deploying AI systems within enterprise environments.
- Display strong problem solving skills and a meticulous approach to debugging complex model behaviors.
- Commitment to maintaining high standards of code quality, documentation, and reproducible experimentation.
- Willingness to collaborate in a fast paced environment where priorities may shift based on business needs.
- Adhere to company policies and contribute positively to a diverse, inclusive, and collaborative team culture.
Nice to have
- Direct experience developing AI applications for enterprise, scientific, or technical domains with complex requirements.
- Hands on work with distributed training systems, large scale inference optimization, and GPU resource management strategies.
- Background in designing automated evaluation suites, data flywheel implementations, and human feedback driven learning pipelines.
- Familiarity with multimodal models, AI agent frameworks, and tool calling or function calling systems.
- A strong track record of publishing research in top tier AI and NLP conferences or notable open source contributions to the ML community.
Practical notes
This is a full time position based in Shanghai. Candidates must be available to work onsite in accordance with local regulations and team collaboration needs. No specific travel requirements are outlined in the source material. The role involves working with sensitive enterprise data and requires adherence to strict security and compliance standards. Applicants should be prepared for a rigorous evaluation process to assess technical depth and practical problem solving abilities. The position reports to senior engineering leadership and requires immediate readiness to contribute at a high level.