Senior Software Engineer, Backend
Job description
About the role
You will design and operate the stateful, low-latency runtime that powers voice and chat AI agents, handling LLM streaming and conversation state management with resilience across multiple channels. You will own the distributed session management layer, implementing affinity, checkpointing, crash recovery, and consistency under concurrent access to ensure reliable agent behavior. You will build a serverless-style function execution platform where customers can deploy custom logic, including orchestration, container lifecycle, autoscaling, and versioned rollout capabilities. You will craft developer experience through CLI tools, local development environments, and test execution frameworks that accelerate iteration and confidence in shipping. You will enforce a high bar for production quality by driving observability, structured incident response, and engineering best practices across the backend team. You will actively use AI-native workflows, leveraging LLMs and AI-assisted tools in your daily development to solve problems that would otherwise be impractical. You will collaborate closely with product and AI teams to translate real-time customer interaction requirements into robust backend services. You will take end-to-end ownership of features from design through production, debugging live issues and refining systems for scale and efficiency.
Key facts
What you'll do
- Architect and maintain a real-time inference runtime for voice and chat AI agents with sub-second latency targets.
- Implement session state management that supports affinity, checkpointing, and crash recovery for long-running conversations.
- Design a serverless function execution platform with container lifecycle, autoscaling, and secure isolation boundaries.
- Build versioned deployment workflows and progressive rollout mechanisms for customer-defined logic.
- Develop CLI and local simulation tools that mirror production behavior to accelerate debugging and iteration.
- Create comprehensive test infrastructure including load, fault injection, and chaos experiments for backend services.
- Define and drive observability standards, including metrics, traces, and logs for distributed agent workflows.
- Establish incident response playbooks and own on-call rotations for critical production services.
- Partner with AI researchers to optimize data paths for streaming token generation and model orchestration.
- Refine developer workflows by reducing build, test, and deploy cycles through automation and caching strategies.
- Evaluate and integrate workflow engines to coordinate long-running, durable processes across services.
- Extend multi-channel support so agents can seamlessly transition between chat, voice, and messaging interfaces.
- Document system behaviors and interfaces to support scaling of engineering teams and new product initiatives.
- Mentor engineers on coding standards, design reviews, and operational best practices for backend systems.
Requirements
- Bring 5+ years of software engineering experience with a strong focus on infrastructure, platform, or systems work.
- Demonstrate deep expertise in both Python and Go, using them as core languages for backend services and tooling.
- Show mastery of distributed systems concepts including consistency, fault tolerance, state management, and concurrency control.
- Have hands-on experience with Kubernetes and cloud-native infrastructure for deploying and operating scalable services.
- Proven track record of building developer-facing tooling such as CLIs, SDKs, and local development environments.
- Communicate clearly in writing and discussion to drive technical decisions and create understandable design documents.
- Maintain a high standard for code quality through thorough testing, thoughtful code review, and sustainable engineering practices.
- Operate comfortably in a production environment with on-call responsibilities, incident response, and direct ownership of service reliability.
- Apply an AI-native mindset, actively using LLMs and AI-assisted tools to enhance development speed and solve complex backend challenges.
- Thrive in a fast-paced setting where requirements evolve and systems must adapt without sacrificing stability.
- Collaborate effectively with cross-functional teams, aligning backend work with product and research priorities.
- Commit to learning new technologies and patterns as the platform and product roadmap expand over time.
- Ensure solutions comply with security and privacy best practices when handling customer conversation data.
- Balance trade-offs between performance, cost, and implementation complexity in distributed system design.
Nice to have
- Experience with real-time voice or streaming media systems that handle continuous data flows.
- Hands-on work with LLM integration, including streaming inference, prompt orchestration, and retrieval-augmented generation.
- Background building serverless or function-as-a-service platforms with elastic scaling.
- Familiarity with workflow engines like Temporal, Argo, or Airflow for long-running, durable processes.
- Prior involvement in conversational AI or speech technology domains and related tooling.
Practical notes
This is a full-time remote position based in Canada. The engagement is full-time with expectations for synchronous collaboration during standard business hours. No specific travel or visa requirements are outlined at this time, and candidates should be prepared to work effectively in a fully remote setup.