
Project Perseus | Speech & Voice AI Analyst
Job description
Project Perseus | Speech & Voice AI Analyst at Welo Global.
About the role
You will manage the full lifecycle of data preparation for speech and voice AI systems, taking ownership of dataset quality from raw audio ingestion to final labeled delivery. You will execute repetitive labeling tasks with rigorous consistency, ensuring that every annotation meets the strict standards required for production model training. This role demands that you exercise sound judgment when resolving ambiguous or edge-case data, translating unclear instructions into precise, actionable outcomes. You will collaborate closely with engineers and researchers to align your work with evolving project requirements and data priorities. You are responsible for maintaining detailed documentation of your labeling decisions and workflows to support reproducibility and auditability. The position requires you to safeguard data integrity by catching errors early and escalating issues before they propagate into model outputs. You will contribute directly to the reliability and performance of AI systems by ensuring that the underlying datasets accurately reflect real-world speech and language patterns. Through continuous attention to detail, you help reduce rework and improve the efficiency of downstream modeling efforts.
Key facts
What you'll do
Perform high-volume labeling of speech and voice data, including transcription, phonetic annotation, and prosodic marking to support model training pipelines.
Conduct rigorous quality checks on labeled datasets, identifying transcription errors, timing mismatches, and inconsistent tags that could degrade model performance.
Follow detailed annotation guidelines and adapt to updated procedures, ensuring strict adherence to project standards across large batches of audio files.
Audit existing labeled data to detect systematic issues or drift, and implement corrective actions to preserve long-term dataset reliability.
Organize and maintain structured metadata for each audio file, including speaker attributes, recording conditions, and channel information to support traceability.
Collaborate with data scientists to resolve ambiguous labeling cases, documenting decisions and edge cases to maintain clarity for future iterations.
Monitor labeling throughput and accuracy metrics, adjusting workflows to balance speed with precision while maintaining high standards of quality.
Support the creation of evaluation benchmarks by preparing clean, well-annotated subsets of data used for model validation and testing.
Track recurring data issues and communicate patterns to engineering and product teams, enabling improvements in data collection and interface design.
Maintain strict version control over labeling schemas and reference materials, ensuring that all team members work from the most current guidelines.
Participate in periodic data reviews where labeled outputs are assessed against original audio to validate alignment and completeness.
Contribute to the development of internal tooling and shortcuts that streamline repetitive labeling tasks without compromising accuracy or consistency.
Assist in the onboarding of new labelers by providing practical examples and guidance derived from hands-on experience with complex audio content.
Act as a gatekeeper for dataset quality, preventing flawed data from advancing to later stages of model development and deployment.
Requirements
Must be authorized to work in the United States without requiring visa sponsorship, as the role is fully onsite and does not support remote arrangements.
Must be currently located in or able to physically commute to one of the following cities: NYC, Seattle, Bellevue, Redmond, San Francisco, Sunnyvale, Burlingame, Austin, Los Angeles, Washington DC, Chicago, Boston.
Must be available to work full-time, 40 hours per week, with consistent scheduling to meet production timelines and data delivery commitments.
Must possess strong attention to detail and the ability to maintain accuracy across repetitive, high-volume labeling tasks over extended periods.
Must be comfortable working directly with audio data, including variations in accent, background noise, and recording quality.
Must be able to follow complex annotation guidelines and apply them consistently across diverse speech and language samples.
Must communicate clearly in writing and verbally, enabling effective collaboration with technical teams and non-technical stakeholders.
Must demonstrate reliability and accountability, ensuring that deliverables are completed on time and to a high standard without constant supervision.
Nice to have
Previous experience in speech recognition, transcription, or phonetic analysis, providing familiarity with common terminology and challenges in the domain.
Familiarity with audio editing tools or data preprocessing workflows that help validate and clean raw recordings before labeling.
Experience working with structured annotation platforms or tools used for large-scale data labeling in production environments.
Basic understanding of machine learning pipelines and how labeled data influences model behavior, including concepts like training, validation, and test splits.
Exposure to multiple languages or dialects, supporting efforts to build inclusive and representative speech and voice models.
Practical notes
This is a 100% onsite position based in California, with no option for remote or hybrid work.
Eligible candidates must be able to commute to one of the specified cities and must already be authorized to work in the United States.
The work schedule is full-time, 40 hours per week, aligned with standard business hours to support coordinated data production cycles.
Candidates should be prepared for repetitive workflows that require patience, focus, and consistent execution over long periods.
The role involves direct interaction with sensitive audio data, requiring professionalism and strict adherence to data handling best practices.
Performance is evaluated primarily through the accuracy and consistency of labeled outputs, as well as responsiveness during collaboration with technical teams.