Senior AI Engineer - LLM Agents
Job description
Senior AI Engineer - LLM Agents at Patsnap.
About the role
The Materials team at Patsnap creates AI tools for R&D professionals to analyze patent and materials science data. You will serve as the primary engineer for our agentic infrastructure, overseeing the development of LLM-based agents and the evaluation systems that ensure our performance exceeds standard models. You will architect and deploy agent systems capable of multi-step reasoning, tool orchestration, and safety guardrails for scientific Q and A and data extraction. In this capacity, you will create and manage memory architectures, MCP servers, and specific agent capabilities within a multi-agent framework. You will collaborate with subject matter experts to design evaluation pipelines that quantify retrieval accuracy and response quality. Furthermore, you will manage the performance, reliability, and observability of production AI services. You will also consult with internal teams on system design, identifying technical constraints and feasibility during product planning. Your work will directly define the reliability and capabilities of the agent stack that powers our commercial products.
Key facts
What you'll do
- Architect and deploy agent systems capable of multi-step reasoning, tool orchestration, and safety guardrails for scientific Q and A and data extraction.
- Create and manage memory architectures, MCP servers, and specific agent capabilities within a multi-agent framework.
- Collaborate with subject matter experts to design evaluation pipelines that quantify retrieval accuracy and response quality.
- Manage the performance, reliability, and observability of production AI services.
- Consult with internal teams on system design, identifying technical constraints and feasibility during product planning.
- Design, implement, and iterate on evaluation sets and metrics to measure agent performance against real-world use cases.
- Take ownership of the end to end lifecycle for AI features, from initial prototype to scaled production deployment.
- Instrument agents and services to capture detailed telemetry for debugging and performance analysis.
- Partner closely with R&D users to understand their workflows and translate requirements into robust agent behaviors.
- Drive the adoption of best practices for prompt engineering, tool use, and chain of thought reasoning across the team.
- Implement retrieval and tool calling patterns that enable agents to access and synthesize complex technical information accurately.
- Continuously refine agent outputs based on evaluation findings, ensuring high standards of factual correctness and safety.
- Maintain and evolve the infrastructure that supports MCP servers and agent tool ecosystems.
- Document system designs, trade offs, and operational procedures to support long term maintainability.
Requirements
- Bachelor degree or higher in computer science, engineering, or a quantitative or physical science field, or equivalent professional experience.
- Minimum 5 years of experience in software or machine learning engineering, with at least 2 years specifically focused on production grade LLM systems.
- Proven experience designing evaluation sets and metrics for LLM agents using tools like Braintrust, LangSmith, promptfoo, DeepEval, or custom frameworks.
- Demonstrated ability to monitor and debug live AI applications using platforms such as Langfuse, Arize Phoenix, OpenTelemetry, or Datadog.
- Proficiency in Python with the ability to independently build and launch services.
- Strong understanding of large language model architectures, training paradigms, and inference patterns.
- Experience building reliable agent workflows, including tool integration, state management, and error handling.
- Familiarity with software engineering best practices such as version control, testing, and continuous integration.
Nice to have
- Experience with RAG pipelines, including vector databases, knowledge graphs, hybrid retrieval, and reranking.
- Background in the Model Context Protocol or agent tool ecosystems.
- Familiarity with chemistry, materials science, or intellectual property data.
- Skills in extracting structured data from complex technical documents like specifications or chemical compositions.
Practical notes
This role offers full ownership of a commercialized agent stack. You will work within a small senior team with direct access to R&D users and domain experts, and your evaluation results will directly influence the product roadmap. The position is based in Singapore and requires full time engagement. There are no specified working hour constraints, travel requirements, or visa conditions listed for this role.