Software Engineer III, Voice AI
NateraRemote (USA)4d ago
AIEngineeringremotecurated-jd
Job description
Software Engineer III, Voice AI at Natera.
About the role
Natera is looking for a Software Engineer III to build and maintain its Voice AI platform. This system handles thousands of patient calls daily, providing automated test status, identity verification, billing support, and intelligent routing. The role focuses on real-time conversational AI systems for Natera's automated patient call center.
Key facts
What you'll do
- Own voice AI features and components through the entire development cycle.
- Participate in design discussions, code reviews, and promote best practices within the Voice AI team.
- Guide technical decisions for voice pipeline optimization, including VAD tuning, turn-taking, interruption handling, and latency management.
- Plan and prioritize tasks in an Agile environment to ensure timely, high-quality delivery.
- Collaborate with Product Managers and stakeholders to refine requirements and scope technical efforts for conversational AI features.
- Monitor voice platform health metrics like call efficacy, ASR accuracy, and per-segment latency, prioritizing improvements based on data.
- Mentor junior team members, sharing knowledge in voice AI architecture, TypeScript, and real-time systems.
- Foster a culture of continuous learning and technical excellence through pair programming and design reviews.
- Document voice AI patterns, integration contracts, and operational runbooks for the team.
- Partner with Product Managers, QA, and clinical operations to gather requirements, validate conversational designs, and guide projects from start to deployment.
- Coordinate with other engineering teams to integrate voice agents with internal services via authenticated APIs.
- Work with the analytics team to ensure voice metrics flow correctly through the data pipeline for reporting and optimization.
- Improve the multi-agent orchestration approach, including tool calling patterns, agent handoff logic, and state management across conversation turns.
- Advocate for high-quality standards and automated testing strategies for conversational AI systems, including voice-specific test patterns.
- Identify and resolve voice-specific UX issues such as ASR errors on medical terminology, silence detection tuning, barge-in recovery, and end-to-end response latency.
Requirements
- 5+ years of overall software development experience, with a focus on scalable backend services.
- 1+ year of experience with voice AI, conversational AI, or real-time audio systems in production environments.
- Hands-on experience with agentic LLM architectures, including tool calling, multi-agent orchestration, prompt engineering, and conversation state management.
- Familiarity with voice AI pipeline components: STT (Deepgram, Azure Speech, OpenAI Whisper), TTS (ElevenLabs, OpenAI, Cartesia), and LLM APIs (OpenAI Realtime API, Anthropic Claude).
- Experience with telephony systems such as Twilio (media streams, SIP, IVR) or similar WebSocket-based audio streaming platforms.
- Understanding of voice-specific challenges: VAD configuration, turn-taking, interruption handling, latency budgets, and audio codec management (mulaw/PCM).
- Solid understanding of the software development lifecycle (SDLC), including build, configuration, release, and deployment.
- Knowledge of microservice architecture and distributed systems best practices.
- Proficiency with AWS services (ECS Fargate, Lambda, DynamoDB, S3, Kafka/MSK, API Gateway).
- Experience with event-driven architecture and message processing (e.g., Apache Kafka, SQS).
- Strong relational database skills (MySQL) and exposure to NoSQL databases (DynamoDB, Redis).
- Demonstrated teamwork skills and a collaborative mindset.
- Excellent communication and organizational skills.
Nice to have
- Experience with RAG architectures (AWS Bedrock, vector stores, embedding models).
Skills & tools
- Node.js/TypeScript (NestJS, Express, async/await, streaming patterns)
- Voice AI Pipeline (telephony ingress, STT, LLM processing, TTS, audio egress)
- Agentic Architecture (multi-agent systems, tool calling, agent handoffs, conversation memory)
- Telephony (Twilio media streams, WebSocket audio streaming, SIP, IVR routing, call recording)
- LLM Integration (OpenAI Realtime API, Anthropic Claude, prompt engineering, structured outputs)
- Database Technologies (MySQL, DynamoDB, Redis/ElastiCache)
- AWS (ECS Fargate, Lambda, DynamoDB, S3, Kafka/MSK, API Gateway, Bedrock, CDK)
- Event Streaming (Apache Kafka, SQS)
- Authentication (Okta JWT flows, OAuth2 client credentials, service-to-service auth patterns)
- Containerization (Docker, ECS task definitions, Fargate deployment)
- CI/CD (GitLab or other pipelines)
- Testing & QA (Jest, conversational system testing, transcript validation, simulated call flows)
- System Monitoring & Troubleshooting (Datadog APM, LLM Observability)
- Compliance (HIPAA requirements for voice systems, PHI handling, zero-retention patterns, encrypted storage)
Practical notes
The pay range for this role is $105,700 to $132,100 USD. Actual compensation packages are determined by factors such as skill set, years and depth of experience, certifications, and specific office location. This range may vary in other locations due to cost of labor considerations.