Senior Platform Engineer
Job description
About the role
You will own the full lifecycle of StackOne's infrastructure, from the code that builds it to the systems that keep it running in production at scale. You will design and manage the pipelines that ship features safely and quickly across both our cloud and our customers' environments. You will partner directly with the CTO and security leadership to define the platform standards that every new service follows from day one. You will implement and maintain the semantic tool discovery and authentication layers that make agentic workflows efficient and secure. You will also be responsible for packaging and supporting our BYOC model so customers can run StackOne inside their own AWS, GCP, or Azure accounts. This role blends hands-on implementation with strategic ownership of reliability, security, and developer experience. You will leverage AI tools where appropriate to accelerate IaC creation, test generation, and runbook drafting while maintaining strict guardrails.
Key facts
What you'll do
Design and operate the AWS estate, including ECS Fargate, Aurora, ElastiCache, MSK, OpenSearch, Lambda, KMS, and related services to ensure reliability, observability, and cost awareness.
Lead the migration of our infrastructure from AWS CDK toward Terraform while maintaining stability and enabling reusable modules for internal and external consumption.
Build and evolve our monorepo CI/CD pipeline, focusing on caching strategies and affected-only incremental test execution to keep feedback loops fast as the codebase scales.
Create and maintain the Terraform modules, container images, runbooks, and documentation required for secure and efficient self-hosted deployments in customers' clouds.
Own the release and upgrade path for self-hosted customers, implementing versioned, signed releases, a supported-version policy, and usage telemetry that respects privacy.
Establish and enforce repository standards by collaborating with security engineers and tech leads to provide secure, deployable templates that accelerate new project onboarding.
Define and drive SLOs, observability dashboards, and incident response processes to improve system resilience and make overnight failures manageable and routine.
Treat infrastructure as a product by building self-service paved roads and tooling that allow product engineers to provision and deploy without waiting on platform teams.
Incorporate AI agents into the infrastructure workflow, using LLMs to generate IaC, scaffold tests, and draft runbooks while ensuring robust security and correctness controls.
Requirements
You must bring 4 or more years of hands-on platform, infrastructure, SRE, or DevOps experience with large-scale AWS environments.
You need deep expertise in Infrastructure as Code, with strong proficiency in Terraform and hands-on experience with AWS CDK or CloudFormation for managing production workloads.
You should have a proven track record of building and optimizing CI/CD and monorepo workflows, including caching and incremental testing techniques that reduce cycle times.
You must be a strong coder in at least one of TypeScript, Python, or Go, capable of building internal tools and automation rather than only editing configuration files.
You must have production experience with containers and Kubernetes, including cluster setup, upgrades, resource limits, and deep troubleshooting of pod lifecycle issues.
You should approach security proactively, embedding secure defaults into pipelines and repository templates and collaborating closely with security teams instead of working around them.
You must be a clear and concise writer, able to create documentation and runbooks that both internal engineers and external customers can follow and trust.
You should operate with end-to-end ownership, scoping work, shipping solutions, measuring outcomes, and automating manual processes instead of relying on checklists.
You must be comfortable working in a fast-moving startup environment where priorities can shift quickly and ownership is expected at all times.
Nice to have
Experience shipping software into customers' own clouds or on-premises environments using BYOC or platform products such as Nuon, Replicated, or Omnistrate.
Hands-on multi-cloud IaC across GCP or Azure, and exposure to distributed systems technologies like Temporal, Kafka or MSK, ClickHouse, or OpenSearch.
Practical notes
This role is full-time based in London.
There may be travel requirements, visa sponsorship considerations, and deadlines for submission of application materials as determined by hiring management.