Staff Back End Engineer, Evals
Job description
About the role
You will architect and own the end-to-end evaluation platform for Hazel, the AI engine for wealth management at Altruist. This role demands deep ownership of the systems that measure factual accuracy, hallucination rates, and compliance across tax planning, financial planning, and investment workflows. You will translate fiduciary-grade requirements into automated quality signals that serve as gatekeepers for every model deployment. You will work directly with backend engineers, product managers, and a network of practicing Certified Financial Planners, CPAs, and tax professionals. Your daily work will involve designing observability dashboards, building golden datasets, and creating verification agents that ensure reliability in a regulated industry. You will establish the CI/CD integration and regression suites that make evaluation a first-class part of our deployment lifecycle. This is an opportunity to define industry standards for AI quality in financial services from the ground up.
Key facts
What you'll do
Architect Hazel's evaluation platform from first principles, including observability, scoring mechanisms, golden datasets, verification agents, and CI/CD integration that define quality standards.
Build production observability and monitoring for AI quality, tracking hallucination rates, factual accuracy, refusal behavior, latency, cost, and domain-specific quality signals across tax planning, financial planning, investment analysis, and operational AI workflows.
Design and implement data curation pipelines that convert real advisor interactions into evaluation datasets, incorporating rigorous sampling strategies, labeling protocols, dataset versioning, and privacy and consent controls required for regulated finance.
Create and steward Hazel's golden datasets in close collaboration with subject matter experts and a network of practicing advisors, CFPs, and tax professionals, translating their tacit expertise into precise, measurable evaluation criteria.
Develop LLM verification agents capable of detecting hallucinations, computational errors, and compliance violations before recommendations ever reach an advisor or client.
Integrate evaluation suites into the deployment pipeline so that every prompt change, model swap, harness modification, or RAG pipeline adjustment runs against regression and acceptance criteria before shipping, making evaluation a mandatory deployment gate.
Partner with the team building Hazel's model-agnostic orchestration harness to evaluate cross-model and cross-provider performance, surface tradeoffs, and inform intelligent routing decisions across Anthropic, OpenAI, and self-hosted models.
Define quality SLOs for each evaluation metric and establish alerting mechanisms that trigger remediation workflows when performance degrades.
Build human-in-the-loop review workflows that scale across every surface of Hazel, ensuring that edge cases and high-risk scenarios receive appropriate expert attention.
Create automated regression suites that continuously validate factual accuracy, logical consistency, and regulatory compliance across new features and model updates.
Establish dataset versioning and lineage tracking that allow precise reproduction of evaluation results and support audit requirements.
Implement privacy-preserving evaluation techniques that ensure sensitive advisor and client data are handled in accordance with financial regulations.
Drive experiments to improve evaluation reliability, including prompt engineering, harness design, and ensemble methods that increase confidence in automated judgments.
Collaborate with product teams to align evaluation criteria with evolving product features, ensuring that quality metrics reflect real advisor and client needs.
Requirements
Demonstrate advanced proficiency in backend engineering with a strong command of at least one systems programming language such as Go, Rust, or Python in production environments.
Bring experience designing and operating distributed systems that handle high throughput, low latency, and strict reliability requirements in regulated settings.
Show a proven track record of building observability platforms, including metrics, tracing, and logging, for complex AI or data-intensive services.
Possess deep understanding of evaluation methodologies for large language models, including prompt engineering, harness design, and benchmark construction.
Exhibit familiarity with financial services compliance, data privacy regulations, and the unique constraints of working with regulated data in wealth management.
Display strong collaboration skills, with the ability to work effectively alongside subject matter experts who are not engineers, such as CFPs, CPAs, and tax planners.
Commit to a growth mindset, taking ownership of problems end-to-end and using initiative to close gaps where standards or tooling are missing.
Demonstrate comfort working in a fast-paced, mission-critical environment where accuracy, reliability, and fiduciary responsibility are non-negotiable.
Practical notes
This role is hybrid, with four in-office days per week at our San Francisco FiDi location.