Senior AI Engineer (Evaluations
Job description
About the role
We are looking for an AI Engineer to join the team behind the Canvas Agent. In this role, you will design and implement robust LLM-as-a-judge evaluation pipelines that automatically assess the Canvas Agent's multi-step reasoning, tool usage, and conversational helpfulness. You own the end-to-end evaluation strategy that ensures quality, safety, and accuracy of agentic workflows. You will create the grading rubrics and synthetic datasets that baseline agent behavior across complex educational scenarios. You optimize judge prompts to align automated scoring with high quality human evaluation standards. You will analyze failure modes and hallucination patterns to drive concrete engineering improvements. This role bridges evaluation and product development to enable rapid, confident iteration. Your work will act as the guardian of trust in AI interactions for educators and students.
Key facts
What you'll do
- Design the Evaluation Framework: Build and maintain scalable LLM-as-a-judge pipelines to automatically score the Canvas Agent's actions, responses, and tool usage across a variety of complex educational workflows.
- Develop Rubrics & Datasets: Create comprehensive grading rubrics and curate high-quality "golden" datasets (both real and synthetically generated) to baseline and test the agent's performance.
- Optimize Judge Prompts: Engineer and iterate on prompts for the judges, ensuring automated scoring aligns with high quality evaluations.
- Full-Stack Contribution: Step beyond evaluation pipelines to participate in full-stack product development as needed, collaborating with the team to build and refine the core AI features, UI components, and application architecture.
- Accelerate Iteration: Integrate your automated evaluations directly into our CI/CD pipelines, creating a "paved path" that allows our AI product teams to ship updates with high velocity and total confidence.
- Analyze & Report: Monitor evaluation metrics to identify failure modes, hallucination rates, and regressions. Translate these subjective quality signals into objective, actionable engineering tasks.
- Define and track core KPIs for agent reliability, correctness, and usability within educational contexts.
- Partner with product managers to translate user needs into measurable evaluation criteria and benchmarks.
- Implement experiments to compare model versions and guide decisions on feature rollouts based on evaluation outcomes.
- Contribute to open source evaluation tools and best practices where appropriate to strengthen the broader AI community.
- Collaborate with data science researchers to validate evaluation methodologies and publish findings.
- Build dashboards and alerts that surface evaluation trends and anomalies to stakeholders in real time.
- Support debugging and root cause analysis when agents fail in realistic educational scenarios.
- Maintain documentation for evaluation protocols, datasets, and infrastructure assumptions.
Requirements
- LLM Evaluation Experience: A background that clearly demonstrates hands-on experience designing and writing automated tests for LLMs. You must have a proven track record of writing and implementing LLM-as-a-judge evaluations in real-world scenarios.
- Technical Stack: A strong background in Python is highly preferred, or a demonstrated willingness and ability to learn it quickly. You should also have professional experience in full-stack or backend engineering to support both the evaluation infrastructure and general product development needs (experience with TypeScript/Node.js is a plus).
- Prompt Engineering Expertise: Deep understanding of how to reliably prompt models for classification, extraction, and grading tasks without falling prey to common biases (e.g., position bias, verbosity bias).
- Evaluation Tooling: Familiarity with modern LLM observability and evaluation frameworks (e.g., LangSmith, Braintrust, Ragas, promptfoo, or similar).
- Agentic Architectures: Strong conceptual understanding of how AI agents plan and execute tool calls (e.g., ReAct, tool-use APIs) so you can effectively evaluate multi-step workflows.
- Analytical Mindset: An ability to translate highly subjective concepts (like "helpfulness" or "tone") into rigorous, trackable metrics.
- Collaboration Skills: Ability to work across product and engineering teams to understand the Canvas Agent's use cases and align evaluation criteria with user needs.
- Reliability Focus: A meticulous approach to testing, debugging, and monitoring that ensures robust and reproducible evaluation results over time.
Onsite Collaboration Requirement: This role requires working onsite on Tuesday and Wednesday, with Thursday strongly encouraged as part of our company's in-person collaboration model.
Why Join Us?
In this role, you aren't just testing a feature; you are the guardian of quality for how AI interacts with the world of education. Your work will be the catalyst that allows Instructure to move faster than ever, safely turning static tools into active participants in the learning journey. We value speed, rigorous testing, and scalability. You will have the opportunity to work at the cutting edge of AI evaluations, setting the standards for how we measure the success of complex agentic behavior. If you enjoy building the "guardrails" that make cutting-edge AI reliable enough for real-world classrooms, this is the role for you.
We invest in our employees' growth and success through thoughtful mentorship, hack weeks, internal conferences, and a culture that values ownership, experimentation, and innovation. If you're excited by building foundational AI infrastructure that empowers others and drives real-world impact, we'd love to meet you.
Get in on all the awesome at Instructure!
We offer competitive, meaningful benefits in every country where we operate. While they vary by location, here's a general idea of what you can expect:
- Competitive compensation, plus all full-time benefits.
- Generous paid time off, including vacation and holidays.
- Comprehensive health, dental, and vision coverage.
- Retirement savings options with company match.
- Professional development stipends for learning and growth.
- Access to wellness programs and mental health resources.