Senior DevOps Lead
Job description
About the role
You will lead the design, implementation, operation, and continuous enhancement of LivePerson's enterprise observability platforms centered on logs, metrics, traces, alerting, and synthetic monitoring. You will own and optimize large-scale observability pipelines that process high volumes of telemetry data daily using a modern stack on Google Cloud Platform. You will design, build, and maintain scalable, Kubernetes-based observability services using cloud-native practices and drive standards for hundreds of engineering users. You will partner closely with DevOps, SRE, Engineering, NOC, Security, and vendor teams to deliver reliable and actionable solutions. You will evaluate emerging observability technologies and drive adoption of practices such as OpenTelemetry and distributed tracing. You will lead the design and development of an end-to-end Synthetic Monitoring platform running hundreds of daily tests on GCP Spot VMs to enable proactive service validation. You will provide technical leadership and mentorship while shaping best practices across observability, DevOps, and cloud engineering.
Key facts
What you'll do
Lead the design, implementation, operation, and continuous improvement of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
Own and optimize large-scale observability pipelines processing high volumes of telemetry data daily, leveraging technologies such as Filebeat, Kafka, Logstash, Elastic Cloud, Prometheus, OpenTelemetry, Grafana Labs, Zabbix, Anodot, and related tools.
Design, build, and maintain scalable, Kubernetes-based observability services using Helm, CI/CD pipelines, GCP, GKE, Docker, and cloud-native best practices.
Define and establish observability standards, dashboards, alerting frameworks, best practices, and onboarding resources for hundreds of engineering users.
Partner closely with DevOps, SRE, Engineering, NOC, Security, and vendor teams to deliver reliable, scalable, and actionable observability solutions.
Evaluate emerging observability technologies and drive adoption of modern practices in areas such as OpenTelemetry, distributed tracing, anomaly detection, and proactive monitoring.
Lead the design and development of an end-to-end Synthetic Monitoring platform, running hundreds of daily synthetic tests on GCP Spot VMs to enable proactive service validation and help engineering teams identify potential issues before they impact customers.
Provide technical leadership and mentorship, driving best practices across observability, DevOps, and cloud engineering.
Develop and maintain detailed runbooks, operational playbooks, and incident response procedures for observability services to ensure clarity and consistency during operations.
Collaborate with product and platform teams to ensure observability requirements are considered early in design discussions and to enable fast, data-driven decision-making.
Implement cost-optimization strategies for observability data storage, retention, and processing across cloud and on-premises environments while maintaining performance and reliability.
Champion reliability engineering principles by defining service-level objectives, error budgets, and alerting policies that balance risk and velocity.
Work with security and compliance teams to ensure observability pipelines adhere to data governance, privacy, and regulatory requirements.
Drive continuous improvement through post-incident reviews, blameless investigations, and feedback loops that enhance system resilience and team workflows.
Requirements
5+ years of experience in software engineering, DevOps, SRE, or a related discipline, with a strong background in application development and cloud engineering.
Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
Strong experience with Kubernetes and containerization technologies, including Docker.
Extensive experience with observability and monitoring technologies such as Grafana Labs, Captain Hook, Zabbix, Fluentd, ELK, Kafka, and Prometheus.
Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation.
Experience working with cloud platforms such as GCP, AWS, or Azure, including compute, storage, networking, and related cloud services.
Strong programming and scripting skills in one or more languages, such as Go, Java, JavaScript, or Python.
Experience designing and operating highly available, scalable, production-grade systems.
Strong understanding of CI/CD practices and cloud-native development methodologies.
Experience with OpenTelemetry Collector and Grafana Agent is highly preferred.
Excellent problem-solving, communication, and collaboration skills, with the ability to influence technical direction across engineering teams.
Nice to have
Experience with OpenTelemetry Collector and Grafana Agent is highly preferred.
Practical notes
Location: Sofia, Bulgaria
Remote Status: Fully Remote (#LI-Remote)