Principal AI Engineer
Job description
About the role
You will own the design and delivery of the internal evaluation and iteration infrastructure that turns raw LLM capabilities into reliable, production-grade conversational commerce agents. You will define how thousands of conversations are evaluated automatically, ensuring that every AI behavior aligns with brand quality and customer experience standards. You will work at the intersection of AI engineering, product, and data, translating non-deterministic model outputs into predictable, measurable outcomes for the business. You will lead cross-functional collaboration to close the gap between prototype experiments and scalable, monitorable features. You will mentor engineers and squads on best practices for building and operating AI systems safely. You will accelerate the company's primary KPI by shortening the cycle from idea to calibrated agent in production.
Key facts
1. Architect the Evaluation "Factory"
- End-to-End Platform Ownership: Architect and lead the development of our internal evaluation platform, moving the needle from manual testing to a fully automated lifecycle (from LLM-as-a-judge creation to production monitoring).
- Accelerate Time-to-Market: Directly impact our primary KPI by designing tools and workflows that drastically reduce the time it takes to deliver a calibrated, production-ready agent.
- Infrastructure Collaboration: Partner with the Orchestration team to build the robust, scalable infrastructure required to run complex evals and agentic simulations at scale.
2. Scaling AI Expertise
- Squad Empowerment: Serve as the "AI Technical Lead" for product squads, guiding them through the complexities of agent design, failure analysis, and prompting best practices.
- Decentralize Quality: Instead of being a bottleneck, you will build the "paved road" that allows product squads to become autonomous in measuring and maintaining their own agent quality.
- Standard Setting: Define what "good" looks like for AI at [Company Name]. You'll translate non-deterministic AI behavior into predictable engineering metrics that the whole organization can trust.
3. Engineering Leadership
- Mentor & Level Up: Bridge the gap between traditional software engineering and AI. You'll mentor engineers on how to apply rigorous system design to the world of LLMs and agents.
- Continuous Observability: Take ownership of the feedback loop, ensuring that production insights from our agents directly inform the next iteration of our evaluation datasets.
Requirements
- 8+ Years of Engineering Excellence: You are a Staff-level engineer first. You've built systems that handle high scale, and you know how to architect for long-term maintainability and performance.
- Agentic Curiosity: You've moved beyond the "chatbot" phase and are actively experimenting with AI Agents. You understand that the challenge isn't the prompt, but the orchestration, state management, and reliability of the agent's actions.
- Systems Thinker (Non-Deterministic Mindset): You recognize that AI is probabilistic. You are excited by the challenge of building deterministic "wrappers" and Evaluation loops around models to make them safe for production.
- The "Applied" Edge: You likely come from a background in distributed systems, internal platforms, or developer tooling, and you're now applying that rigor to the AI stack.
- Beyond the Wrapper: You have serious experience moving beyond simple API calls to architecting multi-stage AI orchestrations (agents, chained workflows, or custom runtime logic).
- Orchestration Experience: Even if you aren't an AI researcher, you have experience building complex, multi-step workflows (e.g., temporal systems, state machines, or event-driven architectures) and want to apply this to Agentic loops.
- Reliability Obsession: You understand why "versioning, testing, and monitoring" are non-negotiable when shipping AI features that impact real customers.
- Production First: You care deeply about shipping systems that are observable, debuggable, and maintainable in long-running production environments.
Nice to have
- Experience building internal platforms or developer tools that enable other engineers to ship reliably.
- Background in evaluating language models or running benchmark suites.
- Contributions to open source AI tooling or agentic frameworks.
Practical notes
This role is full time based in Paris. We are an equal opportunity employer and require reliable local presence. No visa sponsorship is available for this role at this time. There are no deadlines to apply, but early submission is strongly encouraged.