Staff AI Platform & Agent Runtime Engineer
Job description
About the role
EQ Bank is seeking a Staff AI Platform & Agent Runtime Engineer to establish the foundational platform for enterprise-scale artificial intelligence and autonomous agent execution. The successful candidate will architect and operate the core systems that power next-generation AI agents, managing their journey from initial experimentation to robust production deployment. This position is responsible for enabling teams throughout the organization to develop, deploy, and scale intelligent agentic workloads with a high degree of security and reliability. You will operate at the convergence of platform engineering, MLOps, and agentic AI, defining the runtime environment, tooling, and developer experience that accelerates the adoption of AI across all business units. This is a hands-on leadership role requiring deep expertise in solving complex infrastructure challenges and setting the long-term technical direction for a rapidly evolving field. You will be the primary architect ensuring that the platform is not only functional but also scalable and maintainable for future demands. Your work will directly influence the speed and safety with which the organization can innovate with AI technologies. You will act as a bridge between cutting-edge AI research and stable, production-ready infrastructure. Ultimately, you will own the systems that ensure AI agents run predictably, securely, and efficiently at scale.
Key Facts
-
Location: Canada
-
Engagement: Full-time.
-
Compensation: Annual salary range is 180,000 to 220,000 USD, based on years of experience within the range of 3 to 5 years.
What you'll do
- Design intake pipelines responsible for accepting agent configurations while implementing rigorous safeguards for sensitive data to prevent unauthorized access or leakage.
- Construct runtime components for agent execution, focusing on reliability and performance using specified products and technologies that meet enterprise standards.
- Conduct thorough reviews of deployment artifacts prior to their promotion into live environments to ensure they meet strict security and quality benchmarks.
- Coordinate shipping procedures across multiple teams to maintain high availability and comprehensive observability for all agent services running in production.
- Forge partnerships with various teams across the organization to ensure platform capabilities align with strategic business objectives and emerging needs.
- Establish secure and reliable pathways for experimental code to progress safely into production clusters without compromising operational integrity.
- Automate validation checks for agent behavior under realistic traffic patterns to identify potential issues before they impact end-users.
- Maintain toolchains and development workflows to allow developers to iterate rapidly without compromising core operational flows or system stability.
- Define and enforce standards for AI execution environments to promote consistency and best practices across the engineering organization.
- Troubleshoot complex runtime issues that span infrastructure, networking, and application logic to ensure minimal downtime and high reliability.
- Implement monitoring and alerting systems that provide deep insights into the performance and health of AI agent workloads around the clock.
- Optimize the developer experience by building intuitive interfaces and tools that abstract complexity while providing powerful capabilities.
- Collaborate with security teams to ensure all platform components comply with industry regulations and internal governance policies.
- Drive the evolution of the platform by researching new technologies and proposing improvements that enhance scalability and efficiency.
- Act as a subject matter expert for AI platform decisions, guiding other engineers and stakeholders on technical direction.
Requirements
- Three years of hands-on experience with agent platforms is mandatory for this role, demonstrating a track record of building or managing such systems.
- Proven skills in operating large language model tooling within production environments are required to ensure models run efficiently and reliably.
- A strong understanding of container orchestration principles is essential, including managing workloads in Kubernetes or similar systems.
- Secure network design paradigms must be understood deeply to implement robust isolation and communication strategies for AI workloads.
- You must be fully comfortable with the management of infrastructure through code, utilizing infrastructure-as-code practices and methodologies consistently.
- Experience with MLOps frameworks and toolchains is necessary to support the full lifecycle of machine learning models and agents.
- A solid grasp of software development best practices, including version control, testing, and continuous integration, is required.
- The ability to work effectively in a fast-paced environment while maintaining a strong focus on documentation and knowledge sharing is mandatory.
Nice to have
No additional requirements or preferences have been specified for this position.
Practical notes
Candidates are advised to apply promptly as the role is full time and based in Toronto. The position requires a commitment to the defined work hours and location. No other practical notes are specified in the source document.