Project Perseus | Data Labeling Associate
Job description
Project Perseus | Data Labeling Associate at Welo Global.
About the role
You will evaluate complex AI model outputs against detailed rubrics to determine quality and accuracy across a wide range of domains. This role requires you to form clear opinions and document the reasoning behind your judgments in a consistent and repeatable manner. You will partner closely with engineers and product teams to surface patterns in model behavior and provide feedback that directly influences system improvements. A significant part of your ownership involves identifying edge cases and ambiguous scenarios where existing guidelines may be unclear or insufficient. You will contribute to maintaining high standards for data integrity by rigorously checking for issues such as logical errors, factual inconsistencies, and safety concerns. This position places you at the center of the evaluation pipeline, where your critical thinking directly impacts the reliability of AI systems in production. You will be expected to communicate your findings clearly and collaborate effectively with cross-functional stakeholders to align on evolving project goals.
Key facts
What you'll do
Investigate intricate model responses to identify subtle inaccuracies, logical flaws, and potential safety risks that automated checks might overlook.
Analyze large volumes of AI generated content to assign quality scores and determine whether outputs meet predefined standards for correctness and usefulness.
Collaborate with technical leads to refine evaluation criteria and propose updates to labeling guidelines based on observed model behaviors.
Perform detailed comparisons between model outputs and reference materials, ensuring that factual claims are verified and properly sourced where applicable.
Detect and categorize hallucinations, inconsistencies, and bias in generated text, then document these findings using structured reporting formats.
Support the calibration of evaluation metrics by conducting manual reviews that validate or challenge automated assessment results.
Engage with multidisciplinary teams to discuss ambiguous cases, challenge assumptions, and reach consensus on the most appropriate labeling decisions.
Monitor model performance trends over time, highlighting recurring failure modes and opportunities for targeted improvements in data quality.
Assist in the creation of validation datasets that reflect real world scenarios, ensuring that evaluation processes remain relevant to actual usage conditions.
Provide actionable feedback to engineering teams, translating qualitative observations into recommendations that can inform model training and deployment strategies.
Qualify edge cases by designing and testing variations of prompts to understand how model behavior changes under different conditions.
Maintain strict data handling protocols to ensure that all reviewed content is treated with confidentiality and processed in compliance with company policies.
Contribute to the development of best practices for evaluation workflows, helping to standardize approaches across multiple projects.
Act as a bridge between technical and non technical stakeholders by clearly articulating evaluation results and the rationale behind assigned scores.
Requirements
Must be authorized to work in the United States without sponsorship, as this role does not support visa sponsorship at any time.
You must be currently located in or able to physically commute to one of the following cities: New York City, Seattle, Bellevue, Redmond, San Francisco, Sunnyvale, Burlingame, Austin, Los Angeles, Washington DC, Chicago, Boston.
You must be available to work full time, 40 hours per week, as this is a high intensity role with strict delivery timelines.
You must possess strong written communication skills, enabling you to produce clear, concise, and well structured evaluation reports.
You must demonstrate the ability to follow detailed guidelines and apply them consistently across diverse tasks and domains.
You must be comfortable forming independent opinions and articulating them professionally during discussions with leads and stakeholders.
You must have a strong attention to detail, with the patience to review complex information and identify subtle issues that others might miss.
You must be willing to engage in critical thinking and challenge model outputs when they do not align with expected standards or real world facts.
Nice to have
Prior experience in AI evaluation, model testing, or data quality assurance is preferred, as this background will help you ramp up quickly on complex tasks.
Familiarity with large language models and an understanding of common failure modes such as hallucination or overconfidence is strongly preferred.
Experience working with structured evaluation frameworks or rubric based scoring systems will be viewed favorably.
Any background in research, journalism, or technical writing that demonstrates rigorous analytical thinking is considered a plus.
Practical notes
This is a 100% onsite position, and remote work is not permitted under any circumstances.
You must be located in or able to commute to one of the designated cities listed in the job details.
The engagement is offered as a 1 year contract, with the possibility of extension based on performance and project needs.
Work authorization must be current and unrestricted for the entire duration of the contract.
All hours are fixed at 40 per week, and adherence to the scheduled work times is essential for successful execution of evaluation cycles.