Senior Software Engineer
Job description
Senior Software Engineer at Thought Machine.
About the role
You will architect and evolve the cloud-agnostic platform infrastructure that powers Thought Machine's global banking core, directly enabling our mission to rid the world's banks of legacy technology. In this role, you will own the design and delivery of infrastructure components that abstract multi-cloud complexity and reduce cognitive load for engineers and bank SREs worldwide. You will build intuitive control planes and automation that turn intricate orchestration, observability, and data streaming into reliable, first-class products for internal and external users. Your work will ensure our platform remains resilient, scalable, and performant at massive scale across AWS, Azure, and GCP. You will play a critical part in advancing observability and optimizing core databases and streaming platforms so our clients receive deep, out-of-the-box monitoring and insights. Ultimately, you will be responsible for deploying and maintaining the foundational systems that make modern, secure, and reliable banking operations possible for our growing global client base.
Key facts
What you'll do
Scale the Core Platform: Evolve, scale, and optimise the cloud-agnostic platform infrastructure powering Vault globally, ensuring it meets the demands of a rapidly expanding financial ecosystem.
Simplify Orchestration: Build intuitive control planes and automation to further abstract away infrastructure complexity, significantly easing the cognitive load for Thought Machine engineers and bank STEs.
Design for Resiliency: Engineer robust, fault-tolerant architectures and Disaster Recovery / Business Continuity Planning (DR/BCP) systems that uphold the highest standards of reliability and availability.
Advance Observability: Enhance our comprehensive observability suites shared across the entire Thought Machine product ecosystem, providing clients with deep, out-of-the-box monitoring and actionable insights.
Optimise Data & Streaming: Scale and optimise core databases and event-streaming platforms to achieve maximum performance and scalability at the lowest possible cost.
Package and Scale Developer-Facing Infrastructure: Create and refine platform utilities that are reliable, secure, and intuitive, directly improving the daily workflows of internal and external developers.
Integrate Third-Party Infrastructure Components: Extend and integrate open-source tools such as Prometheus, OpenTelemetry, and Grafana to seamlessly fit into complex production environments.
Collaborate Cross-Functionally: Work closely with product, engineering, and operations teams to translate business requirements into robust infrastructure solutions that support global banking operations.
Champion Best Practices: Institutionalise patterns, standards, and runbooks that ensure platform components are secure, maintainable, and operate at scale with minimal overhead.
Drive Innovation in Infrastructure: Continuously research and prototype new approaches to infrastructure challenges, ensuring Thought Machine remains at the forefront of cloud-native banking technology.
Contribute to Technical Strategy: Provide technical leadership and guidance on infrastructure roadmaps, influencing long-term decisions that affect scalability, reliability, and developer experience.
Ensure Operational Excellence: Implement monitoring, alerting, and automation that keep our platforms stable, observable, and efficient under varying loads and edge conditions.
Support Global Deployment: Enable infrastructure deployments in multiple regions and cloud environments, maintaining consistency and compliance across all Thought Machine data centers.
Mentor and Enable Teams: Share knowledge and best practices with engineers and SREs, helping them operate complex infrastructure with confidence and reduced friction.
Requirements
Degree in Computer Science, Engineering, or a similar technical field that provides a strong foundation in distributed systems and software design principles.
5+ years of hands-on software engineering experience building and scaling infrastructure platforms, demonstrating a track record of ownership and impact rather than only maintenance.
Strong proficiency in building production-ready platform tooling using Python or Golang, with a focus on clean code, testing, and robust software design.
Deep understanding of container and orchestration internals, including Kubernetes, Docker, and related ecosystem tools, to design solutions that are both powerful and operable.
Hands-on experience extending and integrating open-source tools such as Prometheus, OpenTelemetry, and Grafana into production-grade monitoring and observability pipelines.
Proven ability to package and scale developer-facing infrastructure utilities with a focus on improving system reliability, performance, and overall developer experience.
Ability to work effectively in a fast-paced, collaborative environment where requirements evolve and technical complexity is high.
Commitment to writing high-quality, maintainable code and to participating in code reviews that uphold Thought Machine's engineering standards.
Willingness to engage with challenging technical problems that require deep analysis, creative solutions, and thorough validation before deployment.
Understanding of security and compliance considerations relevant to financial services infrastructure and the implications for platform design.
Capacity to communicate technical concepts clearly to both technical and non-technical stakeholders, ensuring alignment across teams and organizations.
Readiness to travel occasionally for work-related purposes as dictated by project needs and business priorities.
Nice to have
Experience building cloud-agnostic solutions running across AWS, Azure, and GCP, enabling resilient and flexible infrastructure strategies.
Technical depth in databases or event-streaming platforms such as Kafka, Postgres, or DuckDB, with a focus on performance tuning and scalability.
A deep interest in the internals of infrastructure tools, with a track record of troubleshooting, extending, or optimizing them within a professional environment.
Practical notes
25 days holiday and bank holidays.
Flexible working hours.
Cycle-to-work scheme.
Electric car scheme.
Season ticket loan.
Start the day properly with fresh fruit and cereals.
Huge range of healthy (and not-so-healthy) snacks, smoothies and drinks.
Two charity days a year.
Weekly food pop-up.
We actively hire candidates who demonstrate technical excellence in their field and welcome people of all ages and backgrounds, providing everyone with equal access to professional development. You are encouraged to apply even if your experience doesn't accurately match the requirements above.