Project Perseus | Data Labeling Associate
Job description
Project Perseus | Data Labeling Associate at Welo Global.
About the role
You will evaluate complex AI outputs and determine whether they meet high standards of accuracy, coherence, and safety. You will work directly with advanced systems, probing their responses to identify subtle failures that automated tests might miss. Your role involves forming clear opinions about quality and articulating them in structured feedback for technical teams. You will collaborate with leads and stakeholders to discuss edge cases and refine evaluation criteria over time. You will help ensure that model behaviors align with real-world expectations through careful, repeatable assessment. You will contribute to a feedback loop that directly influences how these AI systems are improved and deployed. You will maintain rigorous attention to detail while managing a high volume of tasks in a fast-paced environment. Your work will bridge the gap between raw model performance and trustworthy, human-aligned outcomes.
Key facts
What you'll do
Investigate AI model responses in detail, identifying inconsistencies, logical gaps, and potential safety issues.
Classify outputs using established criteria, assigning clear labels that reflect correctness, completeness, and adherence to policy.
Compare multiple model generations to determine which responses are more accurate, reliable, and contextually appropriate.
Document findings in structured formats, ensuring that observations are traceable, reproducible, and useful for engineering teams.
Participate in labeling sessions where you apply guidelines to real prompts, testing how well instructions are understood and executed.
Flag edge cases and unusual behaviors that could indicate brittleness or unsafe tendencies in model behavior.
Collaborate with leads to clarify labeling standards, resolve ambiguous scenarios, and update procedures when necessary.
Audit previously labeled data to verify consistency and surface patterns where guidelines may have been applied unevenly.
Support the design of small experiments that test model performance under controlled conditions and specific constraints.
Communicate observations and suggestions to stakeholders, translating labeling insights into actionable recommendations.
Assist in maintaining documentation that explains labeling decisions, criteria changes, and observed trends over time.
Monitor workload pacing to ensure that evaluation throughput remains stable while maintaining high accuracy standards.
Provide feedback on tooling and interfaces used for labeling, suggesting improvements that enhance clarity and efficiency.
Contribute to ongoing discussions about what high-quality AI behavior should look like in different domains and use cases.
Requirements
You must be at least 18 years old and legally eligible to work in the United States without sponsorship.
You must reside in or be able to physically commute to one of the following locations: New York City, Seattle, Bellevue, Redmond, San Francisco, Sunnyvale, Burlingame, Austin, Los Angeles, Washington DC, Chicago, Boston.
You must be authorized to work in the United States and not require visa sponsorship for employment.
You must commit to a full-time schedule of 40 hours per week for the duration of the 1-year contract.
You must be available to work on-site for the entire contract period with no remote work options.
You must demonstrate strong critical thinking skills and the ability to form and justify opinions about AI behavior.
You must be comfortable reading, understanding, and applying detailed labeling guidelines consistently.
You must have excellent written communication skills to articulate findings clearly and constructively.
You must be comfortable working independently and as part of a collaborative team evaluating complex systems.
You must be able to manage repetitive tasks while maintaining attention to detail and accuracy.
You must be comfortable with technology and familiar with using software tools to interact with AI systems.
You must have a professional attitude and be willing to receive feedback to improve labeling quality over time.
You must be reliable, punctual, and capable of meeting regular deadlines for labeling assignments.
Nice to have
Prior experience reviewing or evaluating AI systems, language models, or software tools.
Familiarity with basic concepts in machine learning, natural language processing, or AI safety.
Experience contributing to structured evaluation efforts or benchmark-style tasks.
Background in research, analysis, technical writing, or a related field that emphasizes careful judgment.
Exposure to annotation or labeling workflows in previous roles.
Practical notes
This is a 100% onsite position located in California; remote work is not permitted.
You must be able to commute to one of the listed cities and maintain full-time hours during the contract period.
No visa sponsorship is available for this role, and authorization to work in the United States is required.
The engagement is structured as a 1-year contract with potential for extension based on performance and team needs.
Compensation is set at $34 per hour for full-time work.
Candidates should be prepared to discuss their reasoning processes and provide examples of past evaluative work during the application process.