Speech AI Evaluation Specialist
Job description
About the role
The Speech AI Evaluation Specialist Opportunity represents a distinct remote role focused on the Portuguese language ecosystem within Brazil. In this capacity, you will operate as a freelance evaluator, working entirely from a location-independent setup with flexible scheduling parameters. The position is specifically tailored for part-time engagement, requiring a commitment of over 10 hours per week without imposing rigid time constraints. You will directly influence the trajectory of conversational AI systems by assessing their outputs and ensuring alignment with safety and quality protocols. The work involves a deep analysis of spoken Portuguese to determine the efficacy and reliability of AI responses. Your evaluations will serve as a critical checkpoint in the development lifecycle, steering improvements in real-world application. The role demands a high degree of self-discipline and intrinsic motivation, as you will manage your own workflow to meet project standards. Ultimately, you will act as a quality gatekeeper, ensuring that the AI interactions meet the rigorous benchmarks required for advanced language model deployment.
Key facts
What you'll do
- Execute the core tasks and responsibilities associated with the Speech AI Evaluation Specialist position as outlined on the official portal.
- Adhere strictly to the scope of work and guidelines published
- Conduct detailed evaluations of AI-generated Portuguese speech, scrutinizing content for relevance, accuracy, and linguistic appropriateness.
- Analyze conversational flows to assess tone, intent, and overall coherence, determining if they align with desired interaction models.
- Assign numerical scores to each interaction based on established criteria, providing clear and concise rationales for each rating.
- Document the context of every conversation meticulously, capturing nuances that impact the AI's performance and user experience.
- Identify recurring patterns of error or undesirable behavior within the AI outputs to inform iterative improvements in model training.
- Categorize dialogue sequences to track flow and logical progression, ensuring the AI maintains a coherent and user-friendly interaction path.
- Prepare and format evaluation files according to engineering specifications to facilitate seamless integration into the model development pipeline.
- Verify the compliance of every interaction against project guidelines before final sign-off, ensuring data integrity for future use.
- Monitor productivity metrics diligently to ensure the required weekly hour commitment is met without compromising evaluation quality.
- Provide written feedback that is objective and constructive, focusing on specific elements that require adjustment or enhancement.
- Track the clarity and naturalness of speech synthesis or recognition outputs as they pertain to the Portuguese language.
- Contribute to the broader goal of improving AI safety by flagging potentially harmful or misleading conversational outputs.
Requirements
- Possess native-level fluency in Portuguese (Brazilian) to accurately assess linguistic nuances and cultural context.
- Demonstrate a strong command of the Portuguese language, ensuring your evaluations reflect an expert understanding of syntax and semantics.
- Maintain a high level of comprehension in English, rated between B2 and C2, to fully interpret complex written instructions and guidelines.
- Exhibit reliability and a strong work ethic to meet the flexible but demanding hourly commitment of over 10 hours per week.
- Apply strict attention to detail to ensure every dialogue is evaluated thoroughly and consistently against project benchmarks.
- Follow instructions precisely as they are delivered through the official documentation and project portals.
- Utilize analytical skills to dissect conversational data and determine the quality of AI interactions effectively.
- Approach the task with objectivity, separating personal bias from the evaluation criteria to maintain scoring integrity.
Nice to have
The official listing indicates that backgrounds in machine learning tasks, data collection, and preprocessing represent a significant advantage for this role. Direct experience in data evaluation and quality assurance is highlighted as a beneficial qualification. Familiarity with data annotation and labeling activities is also considered a positive factor for candidates entering this position.