
Senior Site Reliability Engineer
Job description
Senior Site Reliability Engineer at Careers.
About the role
Senior Site Reliability Engineer at Careers. <p><strong>Location Details:</strong><strong>&nbsp;</strong></p> <p>At Careers the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.</p> <p><strong>Remote:</strong> This is a remote position, so you'll be working remotely from your home. You may occasionally visit a Careers office to meet with your team for events or meetings. &nbsp;</p> <p><strong>Join our team<br></strong>Our Global Sustaining Engineering team sits at the intersection of software engineering and infrastructure, ensuring the services our customers depend on are fast, resilient, and always available. The listing location is Bulgaria.
Key facts
What you'll do
- <p><strong>Location Details:</strong><strong>&nbsp;</strong></p> <p>At Careers the future of work looks different for each team.
- Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.</p> <p><strong>Remote:</strong> This is a remote position, so you'll be working remotely from your home.
- You may occasionally visit a Careers office to meet with your team for events or meetings.
- &nbsp;</p> <p><strong>Join our team<br></strong>Our Global Sustaining Engineering team sits at the intersection of software engineering and infrastructure, ensuring the services our customers depend on are fast, resilient, and always available.
- As a Senior Site Reliability Engineer, you'll take direct ownership of production services - from initial design through day-to-day operation - while partnering with product, engineering, and security teams to build and maintain business-critical systems.
- In this role, you will deepen your technical expertise and grow your leadership presence by mentoring the next generation of SREs.
- You will also gain hands-on experience with intelligent tooling in real-world workflows.&nbsp;&nbsp;<br>What you'll get to do...<strong><br></strong></p> <ul> <li>Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues</li> <li>Build and maintain alerting pipelines, dashboards, and SLO-driven monitoring strategies using Icinga, Prometheus, and Grafana</li> <li>Lead incident response end-to-end - performing root-cause analysis, authoring blameless post-mortems, and driving corrective actions to closure</li> <li>Develop and extend Infrastructure as Code coverage and build internal tooling that eliminates manual, repetitive operational work</li> <li>Mentor SRE I and SRE II engineers through code reviews, debugging sessions, and knowledge-sharing talks</li> <li>Apply LLM-driven log analysis, anomaly detection, and generative AI tools to accelerate incident response and runbook creation - validating all outputs before use</li> </ul> <p><strong>Your experience should include...<br></strong></p> <ul> <li>6+ years of professional experience in Site Reliability Engineering or Platform Engineering with demonstrated success leading organization-wide reliability programs</li> <li>Deep hands-on expertise with Kubernetes (deployments, operators, custom resources) and Docker in production environments</li> <li>Advanced Linux experience solving problems involving kernel internals, TCP/IP, DNS, and load balancers&nbsp;&nbsp;</li> <li>Proficiency in Python for production-grade automation and scripting, with working knowledge of Bash</li> <li>Expertise in Ansible and at least one additional Infrastructure as Code tool such as Terraform or Pulumi, with hands-on mastery of Icinga, Prometheus, and Grafana</li> <li>Understanding of large language models, embeddings, and basic machine learning pipelines, with the ability to evaluate and integrate AI-ops tools into daily work</li> </ul> <p><strong>You might also have...<br></strong></p> <ul> <li>Experience building and maintaining continuous integration and continuous delivery pipelines using Jenkins, GitLab CI, or GitHub Actions</li> <li>Demonstrated experience defining and managing Service Level Objectives, Service Level Indicators, and Service Level Agreements across production</li> </ul> <p><strong>We've&nbsp;got your back...</strong><strong> </strong>We offer a range of total rewards that may include paid time off, retirement savings (e.g., 401k, pension schemes), bonus/incentive eligibility, equity grants, participation in our employee stock purchase plan, competitive health benefits, and other family-friendly benefits including parental leave.
- Careers's benefits vary based on individual role and location and can be reviewed in more detail during the interview process.<span data-ccp-props="{}">&nbsp;</span><span data-ccp-props="{}">&nbsp;</span></p> <p><em>We encourage you to apply even if your experience or skillset doesn't align perfectly with every requirement.
Requirements
- You will also gain hands-on experience with intelligent tooling in real-world workflows.&nbsp;&nbsp;<br>What you'll get to do...<strong><br></strong></p> <ul> <li>Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues</li> <li>Build and maintain alerting pipelines, dashboards, and SLO-driven monitoring strategies using Icinga, Prometheus, and Grafana</li> <li>Lead incident response end-to-end - performing root-cause analysis, authoring blameless post-mortems, and driving corrective actions to closure</li> <li>Develop and extend Infrastructure as Code coverage and build internal tooling that eliminates manual, repetitive operational work</li> <li>Mentor SRE I and SRE II engineers through code reviews, debugging sessions, and knowledge-sharing talks</li> <li>Apply LLM-driven log analysis, anomaly detection, and generative AI tools to accelerate incident response and runbook creation - validating all outputs before use</li> </ul> <p><strong>Your experience should include...<br></strong></p> <ul> <li>6+ years of professional experience in Site Reliability Engineering or Platform Engineering with demonstrated success leading organization-wide reliability programs</li> <li>Deep hands-on expertise with Kubernetes (deployments, operators, custom resources) and Docker in production environments</li> <li>Advanced Linux experience solving problems involving kernel internals, TCP/IP, DNS, and load balancers&nbsp;&nbsp;</li> <li>Proficiency in Python for production-grade automation and scripting, with working knowledge of Bash</li> <li>Expertise in Ansible and at least one additional Infrastructure as Code tool such as Terraform or Pulumi, with hands-on mastery of Icinga, Prometheus, and Grafana</li> <li>Understanding of large language models, embeddings, and basic machine learning pipelines, with the ability to evaluate and integrate AI-ops tools into daily work</li> </ul> <p><strong>You might also have...<br></strong></p> <ul> <li>Experience building and maintaining continuous integration and continuous delivery pipelines using Jenkins, GitLab CI, or GitHub Actions</li> <li>Demonstrated experience defining and managing Service Level Objectives, Service Level Indicators, and Service Level Agreements across production</li> </ul> <p><strong>We've&nbsp;got your back...</strong><strong> </strong>We offer a range of total rewards that may include paid time off, retirement savings (e.g., 401k, pension schemes), bonus/incentive eligibility, equity grants, participation in our employee stock purchase plan, competitive health benefits, and other family-friendly benefits including parental leave.
- Careers's benefits vary based on individual role and location and can be reviewed in more detail during the interview process.<span data-ccp-props="{}">&nbsp;</span><span data-ccp-props="{}">&nbsp;</span></p> <p><em>We encourage you to apply even if your experience or skillset doesn't align perfectly with every requirement.
About Careers
An American firm registers domains and hosts websites from Arizona offices under Delaware law. Industry rank places it fifth globally with more than sixty two million domains. Core operations focus on small business owners, forming a broad customer base of twenty million. Public markets list this company. High volumes of registered domains reflect widespread use. Consistent service delivery supports entrepreneurs. Many micro and small companies rely on its online tools for visibility.
Software engineering work usually means turning product intent into systems that run in production. Engineers write and review code, break large changes into reviewable pieces, and watch how those changes behave after they ship. Most teams share on-call or incident habits so outages get a clear owner. Collaboration with product and design is part of the job, not an extra. The duties above are the ones this listing actually states for this opening.
About the location
The posting places this role in Bulgaria. Local hiring markets vary in interview pace and compensation bands. Location copy is here so readers understand the geography attached to the title.