Lead Site Reliability Engineer
LuxoftUSA3d ago
EngineeringReliabilityremotecurated-jd
Job description
Lead Site Reliability Engineer at Luxoft.
About the role
This position focuses on maintaining the stability and performance of high-stakes banking systems. You will serve as a senior technical lead, driving SRE standards and collaborating across departments to optimize system health through automation and observability.
Key facts
What you'll do
- Design and maintain reliability practices throughout the software development lifecycle.
- Build and manage automated regression testing frameworks to verify platform stability during updates.
- Implement Infrastructure as Code using Terraform.
- Manage CI/CD pipelines and deployment automation workflows.
- Define and track Service Level Objectives, Service Level Indicators, and error budgets.
- Conduct root cause analysis and manage incident response processes.
- Develop self-healing systems and automated recovery protocols.
- Perform capacity planning and performance tuning for distributed architectures.
Requirements
- Expertise in Azure App Services and Azure networking.
- Proficiency with Azure Monitor and native operational tools.
- Experience with Dynatrace and OpenTelemetry for distributed tracing.
- Ability to manage centralized logging and log aggregation systems.
- Skill in developing alerts and operational dashboards.
- Experience with application lifecycle and release management.
Nice to have
- Background in supporting enterprise-scale applications within regulated industries.
- Proficiency in scripting with Python, PowerShell, or Bash.
- Experience leading incident response teams and post-incident improvement projects.
- Familiarity with Agile and DevOps methodologies.
- Relevant certifications in Azure, Terraform, or SRE disciplines.
- Strong ability to coordinate with cybersecurity, infrastructure, and architecture teams.
Skills & tools
- Azure (Monitor, App Services, Networking)
- Terraform
- Dynatrace
- OpenTelemetry
- CI/CD pipelines
- Python, PowerShell, Bash
- Distributed tracing
- SLO/SLI management
Practical notes
This role requires the ability to lead complex initiatives independently and communicate technical strategies to diverse stakeholders.