Site Reliability Engineer
Job description
About the role
Wheely is currently expanding its infrastructure team following a recent rebuild of its core systems and platform architecture. This specific role is centered on enhancing the overall stability, performance, and security posture of the services that power the business. The successful candidate will own the design and implementation of resilient infrastructure components that guarantee continuous and smooth service delivery for all customers globally. You will operate at the intersection of development and operations, driving initiatives that reduce risk and improve system observability. This position requires a proactive mindset to identify potential failures before they impact users and to own the long-term health of production environments. You will be a key contributor to the technical decisions that shape the reliability and scalability of our cloud-native infrastructure on a daily basis.
Key facts
What you'll do
- Assume ownership for responding to and managing complex system alerts and critical production incidents as they arise.
- Partner closely with development teams to diagnose, isolate, and resolve technical issues across the application stack efficiently.
- Drive the automation of routine operational tasks and the creation of clear, up-to-date technical documentation for runbooks and procedures.
- Actively contribute to the evolution and maintenance of application delivery pipelines to streamline deployment workflows.
- Take part in the planning and execution of the migration of existing infrastructure and application services to new platforms.
- Focus on building and maintaining robust system resilience, ensuring high availability and disaster recovery strategies are effective.
- Design and create dynamic operational environments that are scalable, observable, and easy to manage.
- Evaluate and implement monitoring solutions to provide deep insights into system performance and infrastructure health.
- Collaborate on the development and enforcement of infrastructure standards and best practices across engineering teams.
- Champion the adoption of Infrastructure as Code principles to ensure consistency and repeatability in deployments.
Requirements
- Must demonstrate hands-on, proven experience with major cloud platforms, specifically AWS or GCP, in a production setting.
- Must possess advanced proficiency in Linux troubleshooting, including in-depth knowledge of networking protocols and file system management.
- Must have a solid background in infrastructure automation utilizing configuration management and scripting tools such as Ansible, Terraform, Bash, or Python.
- Must have direct experience with container technologies and orchestration platforms, specifically Docker and Kubernetes, in deployed environments.
- Must hold administrative experience with message brokers, including operational knowledge of RabbitMQ or Kafka.
- Must have administrative experience managing SQL and NoSQL databases, understanding performance tuning and backup strategies.
- Must exhibit strong written and verbal communication skills to facilitate effective team collaboration and to explain technical decisions clearly.
- Must be capable of prioritizing multiple tasks, managing time effectively, and delivering high-quality results against tight deadlines.
- Must show a willingness to collaborate seamlessly with engineers, assist with technical investigations, and actively share knowledge.
- Must possess a demonstrated interest in automating manual processes and a curiosity for exploring and integrating new technologies.
- Must be legally authorized to work in the United Kingdom without sponsorship requirements for this specific role.
- Must have a proven track record of working in fast-paced environments where adaptability and quick learning are essential.
Nice to have
- Experience with specific monitoring and logging tools such as Grafana, Prometheus, Sentry, and Loki.
- Familiarity with CI/CD tools including Jenkins and Github Actions to support pipeline development.
- Previous exposure to job queues and background task processing systems.
Practical notes
This is a full-time position based in an office environment, requiring four days of on-site presence with one designated remote day for focused work. The role offers a competitive salary in the range of £90,000 to £120,000, reflecting the responsibility and impact of the position. Wheely provides a comprehensive benefits package that includes an employee stock options plan, private medical and dental insurance, and life and critical illness cover for you and your family. You will be equipped with a MacBook Pro and a 4k display to support your productivity from day one. Additional perks include a monthly Wheely journey credit for your commute, participation in the Cycle to Work scheme for eco-friendly travel, and a dedicated professional development stipend for continuous learning. The company also offers relocation support, including visa sponsorship and an allowance, to help you transition smoothly to the London team.