Sr. Devops Engineer II
Job description
About the role
DoubleVerify is seeking a Senior DevOps Engineer II to join our team in Paris. In this role, you will be responsible for managing, optimizing, and scaling our complex infrastructure that supports billions of events daily. You will work closely with multiple engineering teams to ensure platform stability, security, and high availability. Your expertise will help us build reliable, scalable, and efficient systems that underpin our core services. This position offers an exciting opportunity to work with advanced cloud and on-premises technologies, contributing to the development of shared services and automation tools that enhance our operational capabilities.
Key facts
What you'll do
- Design, implement, and manage Kubernetes platforms across multiple environments, including Google Cloud Platform (GCP), Amazon Web Services (AWS), and on-premises data centers, supporting over 100 engineering teams.
- Develop and maintain shared services such as Kafka for real-time streaming, Aerospike for NoSQL data storage, and Airflow for workflow orchestration, ensuring these systems operate with over 99.9% uptime and reliability.
- Build automation tools and scripts to streamline platform operations, reduce manual intervention, and improve overall efficiency.
- Enhance and optimize CI/CD pipelines by developing custom operators, deployment management tools, and automation workflows that accelerate software delivery cycles.
- Collaborate with product and engineering teams from the initial planning stages to ensure seamless integration of new features and infrastructure components, minimizing technical debt and future rework.
- Establish and promote best practices in automation, security, observability, and maintainability across the organization, fostering a culture of continuous improvement.
- Lead infrastructure projects in partnership with teams across the US, Israel, and Europe, ensuring cross-regional consistency and performance.
- Use data-driven approaches to identify performance bottlenecks, troubleshoot issues, and implement solutions that improve system reliability and scalability.
- Implement security policies and access controls using tools like Teleport, ensuring secure access management for all infrastructure components.
- Develop and enforce policies using tools such as Kyverno and custom admission controllers to maintain compliance and security standards.
- Monitor system health and performance using observability tools like Prometheus, Mimir, Grafana, and Grafana Alloy, setting up alerts and dashboards for proactive incident management.
- Conduct load testing with K6 to validate system performance under stress and ensure readiness for high-volume traffic scenarios.
- Maintain distributed tracing and logging across the stack to facilitate troubleshooting and root cause analysis.
- Collaborate with global teams to implement best practices in infrastructure management, security, and automation, fostering a unified approach across regions.
Requirements
- A minimum of 5 years of experience in DevOps, Platform Engineering, or Site Reliability Engineering, with a focus on large-scale production infrastructure.
- At least 3 years of hands-on experience managing Kubernetes in production environments, with a preference for experience managing a dedicated Kubernetes platform.
- Strong expertise in cloud platforms, particularly GCP and AWS, with experience working in multi-cloud environments.
- Proficiency in scripting and programming languages such as Python, Go, and Bash, used to develop automation and infrastructure solutions.
- Deep understanding of distributed systems troubleshooting, utilizing observability tools like metrics, logs, and traces to diagnose and resolve issues.
- Experience designing systems that prioritize reliability, scalability, and developer experience.
- Knowledge of container orchestration, service mesh, and infrastructure-as-code practices.
- Ability to work effectively in a collaborative, cross-functional environment with teams in different regions.
- Strong communication skills to articulate complex technical concepts and collaborate with diverse stakeholders.
- Experience working with high-volume, real-time data processing systems is a plus.
Nice to have
- Hands-on experience with Kafka, ArgoCD, Aerospike, or Airflow in production environments.
- Familiarity with service mesh technologies such as Envoy Gateway or Istio.
- Knowledge of GitOps workflows and infrastructure-as-code tools like Terraform or Crossplane.
- Contributions to open-source platform tools or CNCF projects.
- Understanding of Site Reliability Engineering (SRE) principles and practices.
- Experience working in a fast-paced environment with evolving requirements and priorities.
- Knowledge of security best practices in cloud and on-premises environments.
Skills & tools
- Cloud: GCP (primary), AWS
- Orchestration: Kubernetes (GKE, on-premises), ArgoCD, Helm, Kyverno
- Service Mesh: Envoy Gateway
- Infrastructure as Code: Terraform, Crossplane
- Streaming: Kafka
- Databases: Aerospike (NoSQL), various SQL systems
- Workflow: Airflow
- Monitoring: Prometheus, Mimir, Grafana, Grafana Alloy
- Load Testing: K6
- Tracing & Logging: Distributed across the stack, including logs and metrics
- Access Management: Teleport
- Policy Enforcement: Kyverno, custom admission controllers
- CI/CD: GitLab CI/CD, GitHub Actions
Practical notes
This position is based in Paris, and candidates must be eligible to work in France. We offer competitive salaries and comprehensive benefits. The role involves on-site work at our Paris office. Candidates should be prepared to collaborate with global teams across different time zones. We encourage interested applicants to apply promptly, as we are looking to fill this position soon. Candidates with relevant experience and a passion for infrastructure excellence are highly valued. We support a diverse and inclusive workplace and welcome applications from all qualified candidates.