Senior AI Developer
Job description
Senior AI Developer at Hinge Health.
About the role
We are seeking a Senior AI Developer to join our computer vision (CV) Engineering team in Montreal, QC. In this role, you will design, build, and operate the cloud infrastructure, deployment systems, and supporting services that power our next generation of multi-modal models and reasoning agents. This role is focused on productionizing LLMs, VLMs, and multi-modal reasoning systems for millions of active users. You will own the infrastructure and operational foundations for deploying, scaling, validating, monitoring, and continuously improving these systems in production. This includes model serving architectures, agent orchestration services, evaluation and validation pipelines, telemetry, observability, and developer tooling. You will work closely with ML scientists, CV engineers, platform engineers, and SRE partners to translate advanced research systems into secure, reliable, and scalable product capabilities. This role offers an exciting opportunity to help expand the scope of the CV Engineering team into cloud-native deployment of frontier models and reasoning agents.
What you'll do
Deploy and Operate Multi-Modal Models at Scale by building and maintaining production systems for serving LLMs, VLMs, and other multi-modal models with high reliability, low latency, and cost efficiency. Build Validation and Evaluation Systems in collaboration with ML Scientists and QA to create robust offline and online evaluation pipelines for reasoning quality, model behavior, hallucination risk, policy compliance, latency, and regression detection. Validate and Monitor Quality in-the-wild to ensure reasoning systems are working as expected across long tail cases, ensuring that exceptions are caught early before they impact users. Develop Reasoning Agent Infrastructure by partnering with other engineers and SRE to develop orchestration, state management, tool execution, guardrails, and supporting backend services and infrastructure required to run reasoning agents safely and effectively in production. Establish Telemetry and Observability by designing dashboards, traces, logs, alerts, and performance analytics for model inference and agent workflows using modern monitoring platforms. Improve Reliability, Performance, and Cost through optimization of throughput, capacity, fallback behavior, model routing, and infrastructure utilization to reliably support millions of active users. Adopt and Evangelize Best Practices by evaluating emerging tools, frameworks, and deployment patterns for LLM and agent ops, and helping the team standardize on effective practices that scale. Ensure Security and Compliance by implementing security controls, access patterns, and operational safeguards that protect user data and support responsible production use of generative AI systems across regulated environments.
Key facts
Requirements
You must hold a Bachelor's degree in Computer Science, Engineering, or a related field as the baseline academic credential for this role. You bring 3+ years of experience developing and operating cloud-based services, infrastructure, and APIs in production, with demonstrated work on AWS or similar major cloud platforms. You have direct experience deploying and operating production ML systems, especially LLMs, VLMs, or other large-scale model systems, including responsibility for release management, evaluation strategies, and monitoring at scale. You are fluent in observability and operational tooling, including monitoring, logging, tracing, and alerting platforms that provide insight into system health and user impact. You have a strong track record with CI/CD and production deployment workflows across development, staging, and production environments, ensuring safe and repeatable releases. You understand the unique challenges of reasoning systems and can contribute to validation frameworks that assess correctness, safety, and performance under diverse conditions. You communicate effectively with both technical and non-technical stakeholders, translating complex infrastructure constraints into clear product and operational decisions. You work reliably in a fast-paced, cross-functional environment where priorities evolve with product and research directions.
Nice to have
Experience using agentic development workflows, including AI-assisted coding and review (e.g. via Claude), plus reusable skills or agents to improve velocity and quality. Experience with LLM and agent development tooling such as LangSmith, LangChain, LangGraph, or MLflow. Experience with GPU-backed inference systems, model serving optimization, and scaling for latency-sensitive applications. Experience deploying or integrating hosted model APIs such as Anthropic, Gemini, or Bedrock. Experience building validation and telemetry systems for generative AI, including regression testing, quality scoring, and production monitoring. Experience with containerized services and orchestration technologies such as Docker, Kubernetes, ECS, or EKS. Experience with workflow orchestration tools such as Temporal or Step Functions. Experience with Databricks or similar platforms for data, experimentation, evaluation, or ML platform operations. Experience with IAM, secrets management, encryption, and compliance-minded cloud controls. Experience with infrastructure as code, especially Terraform, and agent infrastructure such as orchestration, tool use, execution control, memory/state handling, and guardrails.
Practical notes
Hinge Health maintains a hybrid work model where remote work and in-person work each offer distinct advantages, and the role requires in-office presence three days per week to enable close collaboration and alignment with team rituals. This position is full-time and based in Montreal, QC, and is eligible for standard company benefits as applicable to full-time roles. No compensation details are provided in this job description. The expectations and responsibilities outlined apply to the Senior AI Developer position within the CV Engineering organization focused on productionizing multi-modal models and reasoning agents at scale.