Speech AI Evaluation Specialist
Job description
About the role
This role centers on the evaluation and enhancement of AI-generated Bengali speech and text, focusing on linguistic quality and user experience. You will act as a primary evaluator, assessing how well AI systems perform in natural language understanding and spoken output. The position requires a meticulous approach to identifying issues such as mispronunciation, grammatical errors, and contextual misunderstandings in Bengali. You will contribute directly to improving the safety and reliability of AI applications used by diverse communities. Success in this role depends on your ability to provide structured, consistent, and actionable feedback. This is a critical function in ensuring that AI models meet real-world linguistic standards. Your work will help define best practices for speech evaluation in low-resource language environments.
Key facts
What you'll do
Conduct in-depth evaluations of AI-generated Bengali speech and text using specialized digital platforms.
Analyze conversational outputs for relevance, grammatical integrity, and contextual appropriateness in Bengali.
Apply standardized rating frameworks to assess clarity, fluency, and overall quality of AI responses.
Follow dynamic scenarios and prompts designed to simulate realistic user interactions and edge cases.
Document observations regarding mispronunciations, awkward phrasing, and semantic inaccuracies in detailed reports.
Verify that AI responses adhere to project-specific guidelines and quality benchmarks consistently.
Collaborate implicitly with engineering and linguistic teams through precise feedback on system weaknesses.
Monitor productivity metrics to ensure timely completion of evaluation tasks within assigned sessions.
Identify recurring patterns of error across multiple evaluations to inform iterative model improvements.
Adapt to evolving evaluation criteria and updated project instructions as AI models are refined.
Perform comparative analysis between different AI outputs to determine relative performance levels.
Maintain strict confidentiality regarding all data, conversations, and platform interactions encountered during evaluation.
Demonstrate strong ownership of assigned evaluation batches to ensure comprehensive coverage and accuracy.
Support the broader goal of making AI interactions more transparent and trustworthy for Bengali speakers.
Requirements
You must be a native speaker of Bengali with an intuitive understanding of its dialects and formal registers.
You must possess fluent or advanced proficiency in English, corresponding to levels B2 through C2 on the Common European Framework of reference.
You must be able to work remotely from India, specifically within the Dhaka location context provided.
You must commit to a part-time schedule of at least 10 hours per week on a flexible basis.
You must be available to start the role immediately upon acceptance of the terms.
You must have access to a reliable internet connection and a suitable device for platform access.
You must be comfortable performing repetitive evaluation tasks while maintaining high attention to detail.
You must be able to interpret and apply project guidelines accurately without constant supervision.
Nice to have
Prior experience in machine learning tasks, data collection, or data preprocessing activities is preferred.
Experience in data evaluation, quality assurance, or data annotation and labeling is highly valued.
Understanding of AI model evaluation workflows and performance metrics is considered an advantage.
Familiarity with digital platforms used for crowdsourcing or freelance evaluation is beneficial.
Practical notes
The engagement is structured as freelance or temporary work with no guaranteed ongoing duration.
The schedule requires a minimum of 10 hours per week, with flexibility to choose working hours.
Compensation is set at 4 USD per hour, paid in a timely manner according to project completion.
This position is ideally suited for students, part-time workers, stay-at-home parents, and gig workers.
There may be an undefined duration of work, referred to as TBC, depending on project needs and performance.
You must be located in India, with Dhaka explicitly noted as the operational location for this role.
The role is designed for remote work-from-home arrangements, requiring self-directed time management.
This position represents an opportunity to participate directly in the assessment and refinement of AI technologies that impact how Bengali language interactions are handled. As an evaluator, you will influence the direction of AI safety and usability by identifying critical issues that automated testing might overlook. The role demands intellectual rigor, patience, and a commitment to linguistic precision. It is tailored for individuals seeking supplemental income without sacrificing personal schedule flexibility. Your contributions will support the development of more accurate and culturally relevant AI systems. This is a short-term freelance opportunity that aligns with the needs of diverse professionals and caregivers. The compensation structure is designed to reward consistent quality and productivity. By joining this initiative, you will help establish foundational evaluation practices for Bengali speech AI.