Senior Software Engineer
Job description
About the role
Deepgram is seeking a Senior Software Engineer to join the Model Evaluation & AI Systems team with a primary focus on ensuring the quality and reliability of our advanced voice AI models. In this role, you will own the design and execution of systems that measure, validate, and verify the performance of our speech-to-text, text-to-speech, and emerging generative AI technologies before any release to customers. This position is critical for maintaining the integrity of our product suite and for building the trust our users place in our technology to transcribe and generate voice accurately and reliably. You will act as the technical gatekeeper for model releases, translating complex research metrics into robust, automated validation that prevents regressions and drives product excellence. The work you do will directly influence whether cutting-edge capabilities are delivered safely and consistently to our global customer base. You will be responsible for constructing the evaluation frameworks that provide the evidence stakeholders need to make confident decisions. Ultimately, your contributions will ensure that Deepgram's voice AI remains dependable, performant, and worthy of customer reliance.
Key facts
What you'll do
- Develop and implement rigorous methods for assessing the performance of Deepgram's speech, audio, and AI models across diverse use cases and environments.
- Construct and maintain scalable automated testing workflows capable of processing both batch and streaming data while prioritizing accuracy, reliability, and efficiency.
- Create and operate robust infrastructure for running large-scale evaluations and aggregating results, with the flexibility to utilize GPU resources when necessary for compute-intensive tasks.
- Convert research-defined performance benchmarks and evaluation criteria into concrete, automated checks that models must reliably pass before progressing in the development lifecycle.
- Establish monitoring systems that provide real-time visibility into model performance, enabling the early detection and alerting on quality degradation in live production models.
- Collaborate closely with research, engineering, and product teams to synthesize complex evaluation data into clear, actionable insights that inform release decisions and product roadmaps.
- Integrate quality assurance checks seamlessly into the continuous integration and continuous delivery pipelines to ensure that every change is validated automatically.
- Contribute to the growth and effectiveness of the team by conducting thorough code reviews, facilitating technical discussions, and sharing knowledge to elevate the collective engineering standard.
- Design evaluation frameworks that are modular, maintainable, and extensible, allowing the team to adapt quickly to new model architectures and evaluation requirements.
- Ensure that all testing processes and results are documented clearly to support auditing, reproducibility, and compliance with internal standards.
- Partner with data scientists to translate abstract evaluation metrics into concrete, testable assertions that reflect real-world user expectations and experiences.
- Optimize evaluation pipelines to reduce runtime and resource consumption without sacrificing the depth or accuracy of the analysis performed on voice and AI outputs.
- Implement alerting and dashboarding solutions that provide stakeholders with immediate visibility into regressions, trends, and anomalies across model versions and datasets.
- Act as an expert resource for best practices in model evaluation, testing strategies, and quality engineering within the broader Deepgram organization.
Requirements
- Hold a Bachelor's, Master's, or PhD in Computer Science, AI, Applied Mathematics, or a related field, or possess equivalent practical experience that demonstrates the same level of competence.
- Have at least 5 years of professional experience in software or quality assurance engineering, with a proven history of developing and maintaining testing or evaluation systems at scale.
- Demonstrate strong proficiency in at least one backend or scripting language such as Python, Rust, or Go, with a ability to write clean, efficient, and production-grade code.
- Possess substantial experience in designing, building, and operating automated test pipelines, evaluation frameworks, or complex data processing systems that handle streaming and batch workloads.
- Exhibit sharp analytical abilities and comfort working with quantitative data, including interpreting metrics, defining thresholds, and analyzing statistical variations to draw valid conclusions about model performance.
- Show the capacity to tackle complex technical problems independently, taking ownership of ambiguous challenges and driving them to resolution with minimal supervision.
- Communicate effectively with diverse teams, including engineers, researchers, and product managers, translating technical constraints and findings into clear, collaborative discussions.
- Have a strong attention to detail and a methodical approach to ensuring correctness, reproducibility, and consistency in evaluation results and testing methodologies.
Nice to have
- Practical, hands-on experience evaluating modern AI systems such as large language models, RAG implementations, agentic systems, or multimodal models, including the analysis of their behavior and outputs.
- Familiarity with building tools for broader accessibility, such as using React Native or similar cross-platform mobile frameworks to create interfaces for evaluation or monitoring.
- Direct experience in developing or enhancing evaluation frameworks, benchmarks, or machine learning infrastructure that are used by multiple teams or external contributors.
- A demonstrated commitment to evaluation quality, ensuring correctness, reproducibility, and consistency across all stages of the testing lifecycle.
- Background in voice, audio, speech recognition, or real-time systems, with working knowledge of domain-specific metrics such as WER, MOS, or latency and their implications for user experience.
- Active participation in open-source projects through meaningful contributions, maintenance responsibilities, or community leadership that showcases technical collaboration skills.
- Experience serving as a technical liaison between different teams or platforms, bridging deep architectural understanding with clear communication to align on evaluation strategies.
- Familiarity with cloud infrastructure, containerized environments, and monitoring tools such as Grafana or other systems for tracking anomalies and performance trends over time.
Skills & tools
- Python, Rust, Go
- CI/CD
- Grafana
Practical notes
- Visa sponsorship is not available for this role.
- Travel is not expected as part of the responsibilities for this position.
- The total compensation package includes equity grants and an annual bonus, providing long-term alignment with company success.
- Applications are being accepted now and will be reviewed on a rolling basis until the position is filled.
- This role is fully remote, allowing you to work effectively from any location within the United States while collaborating closely with a distributed engineering team.
- The successful candidate will join a growing team dedicated to building the evaluation infrastructure that protects the quality of Deepgram's AI products.
- You are expected to work standard office hours and participate in synchronous collaboration with teammates across different time zones as needed for planning and review activities.
- This position requires a reliable internet connection and a suitable home working environment to perform the essential duties of the role remotely.