Infrastructure engineer
Job description
About the role
WRITER is building the foundational layer for enterprise superintelligence, and this role is central to ensuring that foundational layer is unbreakable. You will own the systems that guarantee our platform is available, performant, and reliable around the clock for the world's leading enterprises. This position demands a proactive mindset where you solve complex systemic challenges before they impact customers, turning operational theory into resilient practice. You will be instrumental in automating infrastructure across the entire stack, championing reliability best practices, and enabling the product roadmap through robust system design. This is a hybrid role based in our New York City or London hubs, where you will report directly to our director of engineering. Your work will directly determine whether hundreds of companies can deploy and run AI agents grounded in their data without interruption.
Key facts
What you'll do
Investigate and resolve the most complex production incidents by tracing failures to root cause and architecting preventative controls that eliminate recurrence.
Design, implement, and maintain scalable, fault-tolerant infrastructure across AWS, GCP, and Azure, leveraging Kubernetes, Helm, and Terraform to define the platform future.
Automate every manual operational task using Python or Go, systematically removing toil and treating manual on-call work as a defect to be engineered out of the system.
Run agents like Claude Code, Droid, and Codex in daily workflows to draft infrastructure changes, write runbooks, investigate incidents, and review pull requests, building a collective human-agent team.
Own the reliability, performance, and efficiency of WRITER's core services end-to-end, defining SLOs and error budgets and standing behind the outcomes of your work.
Balance immediate critical fixes with long-term platform investments in observability, cost management, and reliability to support multi-year enterprise customer demands.
Collaborate closely with product, security, and engineering peers to provide expert guidance on system design, aligning infrastructure decisions with product revenue and customer impact.
Evaluate and integrate new tools by rigorously challenging the status quo and rejecting solutions that do not fit the problem, ensuring simplicity and operational efficiency.
Encode recurring infrastructure tasks as internal skills that any teammate, human or digital, can execute, compounding team throughput over time.
Lead post-mortems and root-cause analyses to transform incidents into architectural improvements that prevent the same issues from ever happening twice.
Requirements
Bring a track record of 5+ years in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company.
Demonstrate breadth across the stack with experience running systems in production, including Kubernetes, Terraform, Helm, and cloud services on AWS, GCP, and Azure.
Show a commitment to simplicity and via negativa by automating operational tasks and rejecting tools that do not solve the problem at hand.
Prove you can own outcomes end-to-end, including on-call responsibilities, SLO definition, and error budget management for critical services.
Have experience designing scalable, fault-tolerant infrastructure that supports enterprise-grade reliability and performance requirements.
Exhibit strong debugging and incident response skills, with the ability to trace issues to underlying causes and apply learning to prevent recurrence.
Collaborate effectively with cross-functional teams, providing technical leadership and evidence-based guidance on infrastructure decisions.
Thrive in a fast-paced, high-growth environment where you balance strategic platform work with urgent tactical fixes.
Nice to have
Experience with agentic workflows and using AI tools like Claude Code, Droid, and Codex to augment infrastructure engineering tasks.
A background in building and maintaining internal platform engineering tools that enable other teams.
Practical notes
This is a hybrid position based in either our New York City or London hubs.
Please indicate your eligibility to work in the selected country as part of your application.
Candidates must be able to work full-time hours as outlined in the engagement.
Travel is not required for this role.
Visa sponsorship is not available for this role at this time.
Applications will be reviewed on a rolling basis until the role is filled.