Senior Site Reliability Engineer
Job description
About the role
You will define and enforce monitoring and alerting standards that keep Solaris platforms reliable and observable. You will create frameworks that implement key concepts like SLIs, SLOs, and SLAs to guide reliability decisions across product teams. You will champion incident response and reliability best practices to ensure consistent handling of outages and degradation. You will design processes that improve resilience and reduce technical complexity across critical services. You will implement software and tooling to improve resilience and automate operations at scale. You will create generic service and product-specific monitoring dashboards that provide actionable insight. You will evaluate capacity trends and address exhaustion before it impacts service level objectives. You will design systems that automatically failover across different data centers to sustain availability. You will break down high-level architectural goals into small, deliverable tasks for engineering teams. You will evaluate third-party tools and act as the technical point of contact for reliability initiatives.
Key facts
What you'll do
- Architect and standardize monitoring and alerting frameworks that capture service health in real time.
- Establish and maintain SLIs, SLOs, and SLAs that translate business goals into measurable reliability targets.
- Drive incident response standards and runbooks to shorten resolution times and improve post-incident learning.
- Streamline resilience engineering by designing processes that eliminate unnecessary complexity and single points of failure.
- Develop automation scripts and tooling to operate services at scale with minimal manual intervention.
- Build and maintain service dashboards that surface key performance indicators for both internal and external consumers.
- Perform capacity analysis to predict resource exhaustion and prevent breaches against committed service levels.
- Implement cross-data-center failover strategies that ensure continuity during regional outages.
- Translate strategic reliability initiatives into concrete implementation tasks that agile teams can execute.
- Act as a subject matter expert for third-party observability and reliability tools during evaluation and rollout.
- Lead platform changes that affect the entire organization, ensuring safe deployment and rollback strategies.
- Provide deep expertise in financial sector reliability constraints and regulatory considerations.
- Participate in a 24/7 on-call rotation to respond to production incidents at any hour.
- Collaborate closely with product and engineering teams to align reliability work with product roadmaps.
- Continuously assess new tools and techniques to strengthen the overall reliability posture of Solaris.
Requirements
- Hold a degree in Computer Science, Software Engineering, Information Technology or equivalent professional experience.
- Bring 6+ years of experience in DevOps, SRE, or Software Engineering roles within a high-growth environment.
- Demonstrate programming experience in Python, Ruby, Java, or Go to automate and analyze systems.
- Show expertise in Service Discovery and Service Mesh technologies to manage microservice communications.
- Have experience analyzing capacity requirements and addressing exhaustion before it impacts SLOs.
- Proven ability to design systems that support automatic failover across multiple data centers.
- Demonstrate hands-on experience with chaos engineering tools such as Chaos Mesh or Gremlin.
- Experience designing infrastructure patterns that limit the blast radius of any single failure.
- Capability to break down complex architectural goals into manageable, deliverable tasks.
- Comfort evaluating third-party tools and serving as the technical contact for reliability projects.
- Track record of safely rolling out significant platform changes across large organizations.
- Understanding of the unique challenges, compliance needs, and risk profile of the financial sector.
- Willingness to participate in a 24/7 on-call rotation and respond to production incidents.
- Business-level proficiency in English for written and verbal communication; German is a beneficial skill.
- Self-motivated mindset that balances creativity and openness with structured, hands-on execution.
- Agile mindset with a drive to deliver results and continuously reflect on improvements.
Nice to have
No preferred items are specified in the source beyond the stated requirements.
Practical notes
- Engagement is full-time based on the source listing.
- Work location is Berlin, Germany.
- No specific hours, travel, visa, or application deadline information is provided in the source.