Senior / Staff AI Model Engineer
Job description
About the role
You own reliability for an AI copilot embedded within a trading platform serving sophisticated investors. Your responsibility is to design evaluation frameworks, construct benchmarks, and implement quality gates that ensure releases maintain high standards. You collaborate with product and engineering to define precise tool contracts and drive model improvement loops. You also operate monitoring and incident response systems to keep AI behavior predictable and transparent. You will translate ambiguous product goals into concrete technical specifications for AI behavior in a regulated financial environment. Your work directly influences the trustworthiness of automated insights used by professionals managing significant capital. You will partner closely with quant researchers to align model capabilities with evolving trading hypotheses and risk constraints. This role requires balancing rapid iteration with the stringent reliability demanded by institutional users.
Key facts
What you'll do
You will architect evaluation frameworks that assess correctness, safety, execution speed, and the potential for regression in market analysis and trading workflows. You will create curated reference sets, scenario collections, stress cases, and adversarial tests that reflect current market conditions. You will implement automated checkpoints that prevent deployment when critical performance indicators deteriorate. You will work with engineering and product teams to establish clear, deterministic interactions for AI tools, ensuring actions like previews and confirmations are reliable and traceable. You will operate telemetry systems, define alerting policies, conduct root cause analysis, and lead fix-forward initiatives to resolve issues systematically.
You will manage model improvement cycles, including data collection, label quality assurance, offline testing, and controlled rollouts that demonstrate measurable gains in benchmark performance. You will develop metrics and measurement pipelines that translate observations into concrete engineering and product decisions. You will also build a deep understanding of trading mechanics such as margin, short selling, portfolio margin, and risk, and communicate these concepts clearly to users. You will ensure that benchmarks remain relevant as market dynamics shift and new trading strategies emerge. You will mentor other engineers on best practices for testing and deploying AI systems in production environments.
Requirements
You must bring at least eight years of experience shipping production software and demonstrate strong proficiency in at least one programming language. You possess a solid grasp of computer science fundamentals, testing practices, and system architecture that guides complex technical decisions. You have hands-on experience building evaluation frameworks, test harnesses, and benchmark suites for large language models, agents, search systems, retrieval pipelines, ranking models, and recommenders. You have led model improvement cycles involving dataset curation, label quality control, offline experimentation, and deployment that show tangible improvements in benchmark results.
You can design meaningful metrics, build measurement infrastructure, and use data to drive product and engineering choices. You work comfortably across the entire technology stack, debugging model and tool failures, instrumenting services, and collaborating on user experience decisions that enhance safety and trust. You operate with high self-direction and readily move into unfamiliar domains to solve difficult problems. You have a strong sense of ownership for the end-to-end quality of AI features, from data through deployment and monitoring. You communicate complex technical tradeoffs clearly to both technical and non-technical stakeholders.
Nice to have
You have experience with fine-tuning, preference optimization, distillation, or prompt and compiler techniques that improve tool reliability. You have created domain-specific benchmarks and adversarial test suites for high-stakes applications. You have deep experience with trading across multiple asset classes and margin structures.
Skills & tools
Rust, TypeScript, Postgres, React for Web, React Native for Mobile, observability and telemetry tooling, LLM APIs, model serving infrastructure, evaluation and training pipelines.
Practical notes
Please read the About the company section for additional context on Clear Street.
About the company
Clear Street builds institutional-grade execution for sophisticated investors. The platform combines AI-native trading intelligence with low-latency APIs and private market access.
Service teams provide white-glove support across onboarding, integration, and ongoing management. The system is designed for demanding strategies, delivering reliable performance and responsive infrastructure. Clear Street focuses on practical tools that integrate smoothly into existing workflows.