Machine Learning Engineer (Speech/Audio)
Job description
About the role
Plaud creates hardware and software interfaces that transform human conversation into structured intelligence. We are seeking an engineer to join our Global Product R&D Center to refine our audio processing and speech recognition capabilities. In this position, you will own the design and execution of audio data strategies that directly influence product intelligence. You will be responsible for ensuring the reliability and scalability of the speech data lifecycle from raw acquisition to model consumption. The role requires close collaboration with product and research teams to translate business requirements into technical data and model solutions. You will drive the implementation of robust evaluation methodologies that validate the performance of speech and audio systems in real-world scenarios. Your work will ensure that the audio intelligence embedded in Plaud devices meets the highest standards of accuracy and usability. This position is critical to advancing the core speech and audio capabilities that define our product ecosystem.
Key facts
What you'll do
- Architect and maintain large-scale audio and speech data ingestion pipelines, handling collection, cleaning, filtering, labeling, and augmentation across distributed systems.
- Design and implement hotword mining strategies by analyzing ASR outputs and extracting rare or domain-specific triggers with high precision.
- Conduct fine-tuning and evaluation of SpeechLLM and general LLM models to enhance performance on code-switching scenarios, proper noun recognition, and industry-specific terminology.
- Engineer targeted datasets and domain adaptation techniques to optimize models for specific use cases and linguistic environments.
- Develop comprehensive evaluation frameworks and construct test sets to benchmark internal models against commercial and open-source speech and audio solutions.
- Partner with data and research teams to manage TB-scale or tens of thousands of hours of multimodal data streams efficiently and securely.
- Implement monitoring and quality assurance processes to detect data drift, annotation inconsistencies, and model degradation over time.
- Document methodologies, pipelines, and experimental results to ensure reproducibility and knowledge transfer across the engineering organization.
- Lead experiments to validate new speech and audio processing techniques, measuring impact on downstream product features and user experience.
- Contribute to the creation of internal tools and libraries that standardize data handling and model training workflows for speech technologies.
Requirements
- Hold a minimum of 1 year of professional experience in machine learning, speech technology, or large-scale data engineering within a relevant industry or research setting.
- Demonstrate advanced proficiency in Python programming and hands-on experience building models and data pipelines with PyTorch.
- Show expertise in distributed data processing frameworks such as Ray or Spark to handle high-throughput audio workloads.
- Provide evidence of practical background in at least one of the following domains: ASR or SpeechLLM training and fine-tuning, general LLM or ML model training, or the management of large-scale multimodal data pipelines involving TB-scale data or tens of thousands of hours of audio.
- Exhibit strong understanding of audio signal processing, feature extraction, and the challenges associated with real-world speech data.
- Display capability to write clean, modular, and well-tested code that integrates into production-grade machine learning systems.
- Communicate effectively with cross-functional stakeholders to align technical solutions with business objectives and product roadmaps.
- Adhere to best practices in data privacy, security, and compliance relevant to handling conversational audio and personal data.
Nice to have
- Hands-on experience with SpeechLLM architectures or speech self-supervised learning (SSL) concepts, including adaptation and fine-tuning of models like Qwen3-Omni or StepAudio.
- Background in contextual biasing techniques, hotword development, and code-switching ASR improvements to enhance recognition accuracy in multilingual environments.
- Authorship of patents or publications in top-tier academic or applied venues such as ICASSP or Interspeech that demonstrate technical depth in speech and audio processing.
- Prior experience managing data workstreams for large-scale speech projects involving hundreds of thousands of hours of audio and complex annotation schemes.
- Familiarity with end-to-end speech and audio AI workflows, including dataset curation, model training, evaluation, and deployment in production environments.
Practical notes
This is an on-site role based in Singapore. Compensation includes an Employee Stock Ownership Plan (ESOP). Benefits include medical insurance and WICA coverage. Employees receive top-spec laptops, high-performance workstations, and Plaud hardware. The company provides access to frontier AI tools including Claude Code and Gemini. Candidates must be available to work from the Singapore office. No remote or hybrid arrangements are available for this position. There are no specific travel requirements attached to this role. Applicants must meet the listed requirements without exception. The selection process will involve technical assessments and interviews focused on speech, audio, and large-scale data systems.