Software Engineer, Frontier Data Products
Job description
About the role
You will architect and operate the critical systems that manage expert work for AI advancement. In this role, you build the durable backend that captures, coordinates, and validates high-stakes human effort. Your responsibility is to ensure correctness in environments where results can be revised and disagreements occur, often after significant time has passed. You own the layer that connects incoming requests to final, verified outputs. This is infrastructure for stateful, long-duration processes, not simple transactions. You will define the patterns that ensure reliability when human judgment is a required step in the automated flow.
Key facts
What you'll do
Architect multi-stage workflows that coordinate automated processing, expert review, and result reconciliation while handling partial completion and late-stage changes with resilience.
Build orchestration primitives that provide advanced retry mechanisms, failure recovery strategies, idempotency guarantees, and a comprehensive audit trail for every state transition in the system.
Integrate AI model inference into production pipelines while ensuring that human judgment remains actionable and the overall system maintains transparency for debugging complex issues.
Empower cross-functional teams by developing robust APIs and operational tooling that allow Product, Operations, and Machine Learning teams to effectively use, debug, and monitor these systems at scale.
Champion reliability and observability by designing monitoring and alerting for scenarios where a silent failure is more critical than a visible crash, as it can lead to corrupted outputs that are hard to trace.
Define core infrastructure as an early architect for a new strategic area, making foundational technology stack decisions that will guide the company's direction for years to come.
Handle non-deterministic systems by designing solutions that are resilient to the inherent unpredictability of human input and model behavior, especially during long-running, multi-hour jobs that test system stability.
Translate ambiguous and complex systems problems into concrete designs, turning high-level product goals and technical constraints into clean, scalable architectures that can evolve with the business.
Own your services from initial design through deployment, daily operations, and continuous improvement, ensuring that each component meets strict standards for performance and correctness.
Collaborate directly with engineers in Product and Machine Learning to ensure system designs meet real-world needs and that trade-offs are clearly understood and documented.
Debug intricate failures that span automated code, external model APIs, and human reviewers, using deep analytical skills to isolate root causes in distributed environments.
Implement durable state management patterns that keep workflows consistent even when participants are slow, fail, or submit results out of order.
Create interfaces that abstract complexity for downstream consumers, allowing less technical stakeholders to interact with powerful systems without needing to understand every implementation detail.
Continuously evaluate and optimize the cost and latency of operations, balancing the need for thorough oversight with the efficiency of high-volume processing.
Requirements
You have a proven track record of building and maintaining backend services that are reliable and scalable over time, with strong, well-reasoned opinions on technology choices that have stood up under production load.
You can quickly identify the right service boundaries and understand where to apply complexity and where to enforce simplicity, demonstrating an instinct for designing systems that are both powerful and maintainable.
You use async workflows, queues, and other distributed patterns as a core part of your daily practice, and you deeply understand the nuances of idempotency and retries for long-running jobs that cannot be allowed to lose state.
You excel at taking vague product requirements, operational needs, and ML limitations and turning them into a coherent, debuggable system design that anticipates edge cases and failure modes.
You are highly proficient in Python and have hands-on experience with AWS and Postgres, using these tools to build solutions that are secure, performant, and cost-effective in real-world scenarios.
You have used workflow engines like Temporal or similar technologies to manage complex, long-running processes, and you can explain the trade-offs involved in choosing one approach over another.
You write clean, testable code and value observability, ensuring that your systems emit the right metrics, logs, and traces to make diagnosing issues straightforward even for rare edge cases.
You are comfortable making decisions in environments with incomplete information, knowing that the cost of changing architecture later can be high and that early choices must be both thoughtful and adaptable.
Nice to have
Experience contributing to open source projects that involve complex state management or workflow orchestration, demonstrating an ability to collaborate effectively in distributed engineering environments.
Hands-on experience building systems that integrate human-in-the-loop workflows, particularly where correctness, auditability, and timing are critical constraints.
Practical notes
This role requires a commitment to in-person collaboration. You will work from one of our offices five days a week, choosing between San Francisco, New York City, or London.