[Job-30854] AI SDET
Job description
[Job-30854] AI SDET at CI&T.
About the role
The role centers on evaluation and quality for AI-driven enterprise solutions, operating within a globally distributed team. Success depends on rigorous validation of agent behavior and regulatory controls using specialized tooling. You will own the entire lifecycle of quality assurance for AI features, from designing evaluation harnesses to interpreting compliance evidence. This position requires a high degree of autonomy and ownership over complex testing scenarios that do not have precedent. You will act as the primary guardian of correctness, ensuring that AI outputs meet both business and regulatory standards. The work involves close collaboration with product and engineering partners to define what correct behavior looks like. Your ability to translate ambiguous requirements into concrete test strategies will directly determine the reliability of deployed solutions.
Key facts
What you'll do
Automated testing frameworks exercise non-deterministic model outputs, moving beyond traditional QA for AI agents. You will architect and maintain these frameworks to handle variance and randomness inherent in model behavior. The evaluation harness defines and maintains scenarios to validate agent behavior. You will design, implement, and iterate on these scenarios to keep coverage aligned with product risk. Golden paths and adversarial test scenarios challenge AI outputs, ensuring evaluation coverage for complex and edge cases. These test scenarios address non-deterministic model behaviors across defined conditions. Regulatory quality gates such as SR 26-2 and NYDFS Reg 187 are validated through automated checks. You will build and maintain the automated checks that verify solutions meet regulatory standards in targeted environments. This comparison highlights gaps between automated outputs and expected human-level execution. You will analyze these gaps to determine whether failures indicate model error, data issues, or misaligned expectations. Data quality degradation and population segmentation errors are measured across evaluation scenarios. Measurement informs adjustments to reduce bias and improve fairness in model outputs. Human-in-the-loop (HITL) checkpoint reliability is verified within automated evaluation workflows. You will ensure that the integration points between human review and automated systems are robust and trustworthy. Standards enforcement guarantees consistency across user journeys and model responses. You will implement the tests and monitoring that confirm adherence to these standards.
Requirements
The posting states a bachelor's degree requirement. You must possess a bachelor's degree as a minimum educational qualification. Python 3.11+, Pytest, and automated testing frameworks demonstrate advanced proficiency. Proficiency supports the construction and maintenance of test automation. You must have hands-on experience writing and maintaining tests in this stack. Direct experience with AI evaluation tools such as Langfuse, RAGAS, or equivalent patterns is required. Experience with these tools enables effective monitoring and analysis of model behavior. You must have used these specific tools or comparable frameworks in a professional capacity. Model risk management principles and regulatory compliance testing methods are applied rigorously. Application of these methods reduces compliance and operational risk. You must understand and apply these principles in real-world scenarios. Shadow mode deployment strategies incorporate statistical validation approaches. Strategies ensure that comparisons between automated and baseline behavior are meaningful and measurable. You must have experience designing or implementing shadow mode tests. A skeptical, adversarial mindset uncovers edge cases in AI agent behavior. You must actively look for failure modes that standard tests would miss. Uncompromising attention to detail governs the gathering of regulatory evidence. Detail orientation ensures that audit trails and compliance artifacts remain complete and accurate. You must be meticulous in recording and organizing compliance-related data. Collaboration with engineers diagnoses and fixes logic failures in evaluation harnesses. You must work effectively with engineers to debug and improve the testing infrastructure.
Practical notes
CI&T provides inclusion support, health and dental coverage, and learning resources.
Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down unfamiliar problems, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
AI evaluation and testing form a central discipline for ensuring agent reliability and safety.
Python and modern testing frameworks are commonly used for automated quality checks.
Model risk management and regulatory compliance are important in regulated deployments.
Shadow mode testing compares automated outputs against human or baseline behavior.
Cross-functional collaboration with engineers is typical to resolve identified issues.