Jailbreaking Lead, Red Team
Job description
About the role
You personally own the core methodology and tooling that enables Far.Ai to consistently jailbreak the world's leading frontier AI models under real-world constraints. You design and execute multi-stage jailbreak strategies that combine prompt innovation, adversarial training, and exploit synthesis against the most capable systems. You operate at the intersection of research and operations, translating red-team findings into concrete mitigations that model developers can ship. You mentor and elevate other red teamers by codifying your attack patterns into reusable playbooks and training materials. You prioritize targets based on potential impact, balancing severity against exploit reliability and reproducibility across model versions. You maintain rigorous documentation of every bypass, including failure modes, token-level reasoning, and environmental dependencies. You continuously benchmark your techniques against public and private jailbreaks to ensure Far.Ai stays at the absolute frontier of model compromise capabilities.
Key facts
What you'll do
- Break the latest closed and open-weight frontier models on a weekly cadence, producing high-severity, reproducible jailbreaks across chat, reasoning, and agent modalities.
- Architect and iterate on multi-modal jailbreak techniques that exploit training-time weaknesses, deployment-time guardrails, and emergent reasoning flaws.
- Build and maintain a versioned exploit database that tracks which payloads, prompt structures, and model configurations yield consistent bypasses over time.
- Collaborate with the CBRN, cyber, and agent red-team pods to design cross-modal attacks that combine jailbreaking with task-specific goal exploitation.
- Partner with model developers to run tightly scoped pre-deployment tests, providing detailed adversarial reports that inform patch validation and safety guardrail tuning.
- Create standardized measurement frameworks for jailbreak success, defining metrics for severity, generality, and transferability across model families.
- Lead incident response for critical jailbreak findings, coordinating responsible disclosure timelines and producing developer-facing remediation guidance.
- Mentor junior red teamers in advanced search techniques, adversarial prompt engineering, and model interpretability methods that improve attack success rates.
- Represent Far.Ai's red team at top-tier academic and industry venues, publishing novel attack methodologies while responsibly disclosing all technical details.
- Set the technical direction for the jailbreaking practice, defining the roadmap for tooling, automation, and integration with the broader AI safety ecosystem.
Requirements
- You have a proven track record of breaking state-of-the-art language models through jailbreaking, with verifiable results on at least three major frontier model releases.
- You possess deep knowledge of transformer internals, attention mechanisms, and emergent behaviors that can be exploited for unauthorized task execution.
- You are fluent in adversarial attack techniques, including prompt injection, injection via tool use, training-time backdoors, and gradient-based optimization for jailbreak generation.
- You have hands-on experience with red-teaming at scale, managing a portfolio of active exploits across multiple model vendors and deployment configurations.
- You can read and modify model weights and logits at a practical level, using toolchains such as Hugging Face transformers, GPTQ, AWQ, or similar quantization frameworks.
- You are comfortable working in low-resource environments, including regions with restricted API access, and can design jailbreaks that work despite rate limits, logging, and monitoring defenses.
- You have authored clear, reproducible technical reports that enable other teams to replicate your attacks without access to proprietary infrastructure.
- You hold a strong ethical compass and a commitment to safe disclosure, aligning your work with Far.Ai's non-profit mission of advancing AI safety and responsible model development.
Nice to have
- Contributions to open-source red-teaming toolchains or jailbreak datasets that are widely adopted by the security research community.
- Experience advising national-level AI safety or security bodies on evaluation standards for frontier model risks.
- Prior work published at top-tier security or AI conferences that directly informs practical jailbreaking methodology.
Practical notes
This is a full-time role based on a remote-first schedule with no mandatory hours, allowing flexibility across time zones while maintaining overlap for critical collaboration. Travel is rarely required, primarily limited to occasional in-person strategy sessions or lab visits that directly support high-priority red-team initiatives. Candidates must be eligible to work in their country of residence without sponsorship, as Far.Ai cannot provide visa sponsorship for this position. The role reports to the Jailbreaking Lead and maintains a dotted-line relationship with the broader Red Team leadership to ensure alignment with organizational safety objectives and research priorities.