
Senior Site Reliability Engineer
Job description
About the role
The in London owns the reliability and performance of the critical platform that powers private capital markets. You will be responsible for designing and maintaining the infrastructure that ensures continuous service delivery for a global user base operating across multiple time zones. This position requires a calm demeanor and clear communication when navigating complex, high-stakes infrastructure challenges. You will be a key architect in evolving the platform to support the company's ambitious growth trajectory. Daily work involves deep collaboration with application engineers to ensure infrastructure capabilities keep pace with product demands. The role focuses on building systems that are both resilient and efficient at scale. You will drive initiatives that transform operational reliability into a strategic competitive advantage for the business.
Key facts
What you'll do
- Architect and manage compute, storage, and networking components to meet the demands of a global equity platform.
- Streamline production workflows by automating deployment pipelines and eliminating manual intervention points.
- Implement proactive observability strategies using metrics and traces to detect anomalies before they impact users.
- Lead structured incident response efforts, ensuring thorough documentation of root causes and containment measures.
- Develop robust integrations that connect Carta's ecosystem with external financial institutions and data providers.
- Modernize operational tooling by refactoring legacy scripts to enhance maintainability and reduce long-term risk.
- Act as a technical liaison between infrastructure and application teams to align on scalable service designs.
- Validate configurations through rigorous testing in staging environments prior to production deployment.
- Champion the adoption of infrastructure as code principles to enforce consistency and repeatability.
- Optimize system performance by analyzing telemetry data and identifying bottlenecks in the technology stack.
- Guide the implementation of security best practices across all infrastructure components and deployment processes.
- Evaluate emerging technologies and determine their applicability to improving operational efficiency.
- Mentor junior engineers on debugging techniques and reliable system design patterns.
- Support the expansion of Carta's global footprint by ensuring infrastructure readiness for new regions and compliance requirements.
Requirements
- Demonstrate extensive hands-on experience managing cloud infrastructure on AWS platforms.
- Show proficiency in writing Python scripts to automate complex operational tasks and workflows.
- Possess a strong grasp of networking fundamentals, containerization, and microservice communication patterns.
- Have practical experience managing infrastructure through declarative configuration and provisioning tools.
- Be comfortable designing and querying relational databases, with a focus on PostgreSQL systems.
- Exhibit the ability to implement monitoring and alerting solutions using modern observability platforms.
- Provide evidence of successfully managing production incidents and maintaining clear documentation of procedures.
- Display familiarity with leveraging artificial intelligence and large language model tools to enhance daily engineering productivity.
- Hold the right to work and reside in the United Kingdom without sponsorship requirements.
- Commit to working from the London office on a full-time basis as specified by the engagement terms.
Nice to have
Prior experience operating and optimizing CI/CD pipelines is considered a strong advantage for this role.
Skills & Tools
You will utilize a broad spectrum of technologies to ensure the Carta platform remains robust and scalable. Expertise in major cloud providers, particularly AWS, is fundamental. You will work with programming languages such as Python and Java to build automation and services. Infrastructure as Code frameworks will be central to your toolchain, alongside container orchestration platforms like Kubernetes. You will interact with database systems and utilize monitoring platforms to maintain high visibility into system health. Experience with GCP is also valued as the company continues to expand its multi-cloud strategy.
Practical notes
This is a full-time position based in London, England, United Kingdom. Please confirm specific details on the official application page before proceeding.