
Observability DevOps III
Job description
About the role
The Observability DevOps III position at LivePerson represents a leadership opportunity to define and execute the company's observability strategy. As a Principal DevOps Lead, you will own the architecture, reliability, and evolution of logging, metrics, tracing, alerts, and synthetic monitoring across global cloud and on-premises data centers. You will empower thousands of engineers by delivering systems that provide superpowers for development, operations, and platform teams. The role requires deep expertise in modern observability platforms, enterprise-scale tooling, and cloud-native practices on Google Cloud Platform. You will evaluate emerging vendors, standardize best practices, and drive adoption of OpenTelemetry, distributed tracing, and anomaly detection to ensure proactive service validation and system resilience.
What you'll do
Lead the design, implementation, operation, and continuous improvement of LivePerson's observability platforms across logs, metrics, traces, alerts, and synthetic monitoring. Own large-scale observability pipelines processing massive volumes of telemetry data daily, including Filebeat, Kafka, Logstash, ElasticCloud, Prometheus, OpenTelemetry, Grafana Labs, Zabbix, Anodot, and related technologies. Design, build, and optimize scalable Kubernetes-based observability services using Helm charts, CI/CD pipelines, GCP, GKE, Docker, and cloud-native best practices. Define observability standards, dashboards, alerting frameworks, best practices, and onboarding materials for hundreds of engineering users. Collaborate closely with DevOps, SRE, Engineering, NOC, Security, and vendor teams to deliver reliable, scalable, and actionable observability solutions. Evaluate new observability technologies and guide the team in adopting modern practices around OpenTelemetry, distributed tracing, anomaly detection, and proactive monitoring. Design and build the full end-to-end Synthetic Monitoring platform, running hundreds of daily synthetic tests on GCP Spot machines to support engineering teams with proactive service validation. Own the reliability and performance of critical observability backends including Elastic Cloud, Loki, Prometheus, and associated storage layers. Drive automation and self-service tooling to reduce manual overhead and improve time-to-insight for internal customers. Partner with product and platform teams to embed observability into new services from inception through production. Implement cost optimization strategies for high-volume telemetry ingestion, processing, and retention in GCP environments. Champion incident response practices and post-incident reviews to improve system resilience and observability quality.
Requirements
Bachelor's degree in Computer Science, Engineering, or related work experience. 5 years of experience as a software engineer or DevOps engineer, with experience in application development and cloud engineering. Proficient in Kubernetes and containerization technologies (Docker, etc.). Extensive experience with observability tools such as GrafanaLab, CaptainHook, Zabbix, FluentD, ELK, Kafka, and Prometheus. Familiarity with infrastructure as code (IaC) tools like Terraform, Ansible, or CloudFormation. Experience with cloud platforms (AWS, Azure, GCP) and their services related to computing, storage, and networking. Strong programming skills in one or more languages (JavaScript, Java, Go, etc.). The ideal candidate will have experience with OpenTelemetry Collector and Grafana Agent. Ability to work independently and collaboratively in a fast-paced, dynamic environment. Strong written and verbal communication skills for documenting standards and collaborating with cross-functional teams.
Benefits and working conditions
The role is based in Bulgaria and is open to remote collaboration within the specified location. There is no specified engagement type or compensation details provided in the source material. The position requires significant experience with enterprise-scale observability platforms and modern cloud-native tooling. Candidates must be comfortable working with high volumes of telemetry data and complex distributed systems. This position involves both technical leadership and hands-on implementation responsibilities. The successful candidate will influence platform strategy and help shape the future of observability at LivePerson.
Why you'll love working here
LivePerson is a transformational force in how brands and consumers communicate. With over 18,000 brands, including HSBC, Disney, Verizon, and Home Depot, we are on a mission to make life easier for people and brands everywhere through trusted Conversational AI. We believe in a future where conversations are the norm for getting your intentions fulfilled - whatever they are. We are an innovative, intent-driven company that believes in building the future and we are looking for growth minded, unconventional thinkers, developers and builders to join the team. As leaders in enterprise customer conversations, we celebrate diversity, empowering our team to forge impactful conversations globally. LivePerson is a place where uniqueness is embraced, growth is constant, and everyone is empowered to create their own success. We are very proud to have earned recognition from Fast Company, Newsweek, and BuiltIn for being a top innovative, beloved, and remote-friendly workplace.
Belonging at LivePerson
We are committed to fostering a culture of belonging where everyone is empowered to be their authentic selves. At LivePerson, we believe that diverse perspectives drive innovation and create better outcomes for our customers and our communities. We embrace individuality, encourage growth, and provide an environment where all team members can thrive and contribute meaningfully to our mission.