Research Engineer, Real Environments
Job description
About the role
You will work directly with large enterprise clients to capture their operational data and transform it into high-fidelity reinforcement learning environments. These environments are used for capability evaluations and for generating training datasets for frontier AI laboratories. Your primary focus will be on pushing the boundaries of world-building, verifier engineering, and automating the process of creating real-world evaluation systems. This role involves close collaboration with researchers, operators, and AI companies at the forefront of AI infrastructure development. You will be instrumental in building scalable, realistic simulation environments that mirror enterprise workflows and enable AI models to be tested and refined in conditions that closely resemble real-world scenarios. The position offers an exciting opportunity to shape the future of enterprise AI evaluation and contribute to the development of systems that are redefining how AI capabilities are assessed and improved.
Key facts
What you'll do
- Develop, ship, and maintain models for workflow extraction, classification, and grading, ensuring they meet enterprise needs and evaluation standards.
- Engineer autonomous task refinement processes that translate enterprise data preferences into scalable, automated pipelines, reducing manual effort and increasing throughput.
- Deliver high-quality data solutions to enterprise customers and deploy these solutions into live engagements, ensuring seamless integration and operational stability.
- Contribute to defining the future of agentic transformation for enterprises by building evaluation environments that accurately reflect real-world work scenarios.
- Deeply analyze enterprise workflows and operational intricacies to develop comprehensive evaluations that cover all aspects of enterprise work.
- Build end-to-end environments for labs and enterprise clients, including platformizing sandbox applications, loading real enterprise data into these environments, and constructing prompts based on actual workflows.
- Write and implement verifiers that leverage enterprise expertise and golden outputs, ensuring the fidelity and reliability of evaluation environments.
- Systematize the production process of environments to scale throughput while maintaining high standards of quality, realism, and fidelity.
- Collaborate with cross-functional teams to improve infrastructure, automation, and pipeline processes, enabling rapid deployment and iteration of environments.
- Participate in testing, debugging, and refining environments to ensure they are indistinguishable from real-world scenarios, paying close attention to detail.
- Engage with enterprise stakeholders to understand their operational nuances and incorporate their feedback into environment development.
- Contribute to documentation and best practices for environment creation, scaling, and verification processes.
- Stay current with advancements in AI evaluation techniques, simulation fidelity, and enterprise workflows to continually improve environment quality.
- Assist in troubleshooting and resolving issues related to environment deployment, data integration, and verifier accuracy.
- Support research efforts by providing high-quality, realistic environments that enable robust model evaluation and benchmarking.
- Work in a fast-paced, innovative environment where shipping and iteration are prioritized, with a bias toward action.
- Collaborate with the team to identify opportunities for automation and process improvements to increase efficiency and scalability.
- Maintain a focus on quality, realism, and fidelity in all environment development activities to ensure they meet enterprise and research standards.
Requirements
- Proven experience in shipping environments, such as contributing to open-source frameworks, building enterprise environments, or working on agentic evaluations.
- Strong full-stack engineering skills, with the ability to handle infrastructure, application development, and analytics components.
- Demonstrated bias toward action, with a focus on delivering evaluation tools and environments rather than just theoretical discussions.
- Curiosity about model behavior, with a desire to understand and analyze data deeply to improve environment fidelity.
- Exceptional attention to detail to ensure simulations are indistinguishable from real-world scenarios.
- Systems-level thinking skills that enable scaling environment quality without sacrificing realism.
- Experience with automation, pipeline development, and large-scale data processing.
- Ability to work collaboratively in a fast-paced, innovative environment, often involving cross-disciplinary teams.
- Familiarity with enterprise workflows, evaluation methodologies, and data integration techniques.
- Strong problem-solving skills and the ability to troubleshoot complex issues related to environment deployment and verification.
- Excellent communication skills to document processes, share insights, and collaborate effectively with stakeholders.
- Willingness to learn and adapt quickly to new tools, techniques, and enterprise requirements.
- Ability to prioritize tasks effectively and manage multiple projects simultaneously.
Nice to have
- Experience with orchestration or compute services such as Temporal, Modal, or similar platforms.
- Background in synthetic data generation for frontier models, including techniques for creating realistic, high-quality synthetic datasets.
- Past work auditing and scrutinizing industry-standard evaluation benchmarks to ensure accuracy and fairness.
- Familiarity with enterprise IT systems, workflows, and operational data sources.
- Knowledge of simulation fidelity techniques, including physics-based modeling or advanced rendering.
- Experience working with large-scale enterprise data and integrating it into evaluation environments.
Skills & tools
- Full-stack engineering
- Infrastructure development
- Data pipeline automation
- Environment and verifier creation
- Workflow modeling and automation
- Model evaluation and benchmarking
- Cloud computing and orchestration tools
- Data integration and processing
- Simulation and environment design
- Version control and collaborative development practices
Practical notes
This role is based in our San Francisco office and requires on-site presence five days a week. The position offers a competitive salary, benefits, and the opportunity to work at the cutting edge of AI infrastructure development. Candidates should be prepared to engage deeply with enterprise operational data, develop scalable environments, and contribute to high-fidelity simulation systems. The environment demands meticulous attention to detail, a proactive approach to automation, and a strong desire to push the boundaries of what is possible in enterprise AI evaluation. You will be part of a fast-moving, innovative team committed to building systems that are both realistic and scalable, enabling AI models to be evaluated and improved in conditions that closely mirror real-world enterprise workflows.