Senior Software Engineer
Job description
Senior Software Engineer, Platform Engineering ZoomInfo Technologies LLC presents a senior platform engineering opportunity focused on the reliability and delivery capabilities of global data infrastructure. This role centers on owning the full lifecycle of critical platform systems. You will ensure that pipelines remain robust, observable, and efficient for distributed teams. The position demands deep architectural judgment and hands-on execution across multiple cloud and automation technologies. About the role You will serve as the primary owner for the core platform that powers software delivery across the organization. The work involves designing, building, and maintaining the foundational tooling that allows engineering teams to deploy and operate services reliably. You will be responsible for the stability, performance, and security of the automation pipelines that drive the business. This is a hands-on leadership position where you will design systems, implement solutions, and manage them through their entire operational lifespan. You will proactively identify systemic risks and drive long-term improvements that prevent failures before they impact production. Collaboration with security, infrastructure, and product teams is essential to define and implement platform standards. Key facts
-
Location: Canada
-
Engagement: Full-time.
-
Compensation: Base salary range is USD 160,000 to USD 180,000 for this role. What you'll do You will design, operate, and evolve the platform that supports all engineering workflows. A primary focus will be the self-hosted GitHub Actions runner fleet deployed on Google Kubernetes Engine. This includes implementing advanced autoscaling logic, ensuring high reliability, and developing automated cleanup processes for zombie runners. You will architect the GitOps delivery model using ArgoCD, defining the central topology for cluster connectivity. This work involves managing ApplicationSets, ensuring production reliability through patterns like Pod Disruption Budgets and replica configuration, and implementing spot-node avoidance strategies. You will also build event-driven deployment flows by integrating Argo Events with GCP PubSub. A significant portion of the role involves ownership of shared platform components. You will maintain and evolve the Helm charts and ApplicationSet templates that serve every Kubernetes workload. This requires strict versioning, managing release processes, and guaranteeing backward compatibility while improving developer ergonomics. You will lead the modernization of the Jenkins platform, taking ownership of its shared libraries, controllers and agents, plugins, and credential integration. This includes managing JVM upgrades and driving improvements to platform stability. You will build and maintain self-service tooling, creating reusable GitHub Actions workflows, custom actions, and templates. This initiative will help accelerate the migration away from legacy Jenkins pipelines and will involve close partnership with development teams on shared developer tooling. You will provide end-to-end ownership for complex platform projects, guiding them from initial architectural design through implementation and long-term maintenance. This includes identifying system-wide bottlenecks during the design phase to prevent them from becoming production outages. You will own operational excellence end-to-end, which involves managing pipeline reliability from development through production. Your responsibilities will include instrumenting pipelines with Datadog, creating custom dashboards, and setting up monitors for runner log analysis. You will participate in incident response and on-call rotation managed via PagerDuty. This involves designing escalation policies and conducting thorough post-incident reviews. You will also own secrets rotation processes and coordinate vulnerability response activities. Global collaboration is a core part of the position. You will work with distributed teams of architects, infrastructure engineers, and security specialists across different time zones. You will translate pipeline requirements from various product teams into golden paths, reusable workflows, and robust self-service tooling. What you'll bring Your expertise must cover CI/CD platforms with deep experience in Jenkins administration. You should be proficient in Groovy shared-libraries, Docker-based agents, and GitHub Actions self-hosted runners. Mastery of secrets management and pipeline troubleshooting is required. You must possess strong skills in cloud and container infrastructure, specifically Kubernetes administration on GKE. You should understand cluster optimization, networking, Kubernetes internals, and autoscaling behavior. Hands-on experience with GCP and AWS services is necessary. Practical experience with GitOps and delivery tooling is essential. This includes hands-on work with ArgoCD, including ApplicationSets and sync strategies, as well as Helm chart authoring and versioning. You must have experience with Argo Events or comparable event-driven delivery patterns. You must be proficient in writing infrastructure as code using Terraform for GKE clusters, GCP resources, and CI/CD-related modules. You must have hands-on experience operating observability tools to instrument and manage pipelines. You should have a proven track record of creating Datadog dashboards, custom metrics, and monitors. Experience with distributed tracing is required. You must have a background in incident response, including running production on-call, designing escalation policies, and conducting post-incident reviews. You must have practical experience administering container registries, including JFrog Artifactory and GCP Artifact Registry. This includes managing Xray/Curation scanning and image lifecycle policies. You must be comfortable shipping production code in at least one language, such as Go, Python, or Node.js. You should have the ability to build controllers, GCP Cloud Functions, glue services, and other automation. You must have experience or knowledge in the area of LLM applications and AI-assisted developer tooling. This includes building or deploying AI models for engineering workflows or managing related infrastructure. Nice to have No specific additional requirements are outlined for this position. Skills & tools The role requires fluency with the following technologies: Jenkins, Groovy, GitHub Actions, ArgoCD, Helm, Terraform, Datadog, PagerDuty, JFrog Artifactory, GCP, and AWS. Practical notes Please confirm all details regarding schedule, compensation, and specific expectations on the official application page. All facts in this description are sourced