System-driven Senior DevOps/SRE with over 6 years of professional experience scaling high-velocity distributed architectures and cloud-native environments from ~€500M to €1.8B+ annual revenue at AUTODOC. Expert in bridging enterprise-grade cloud infrastructure (GCP/AWS, GKE, Talos OS) with edge topologies and GitOps frameworks (FluxCD, Crossplane, Terraform).
Proven track record of managing hundreds of containerized nodes, cutting infrastructure costs, and executing zero-downtime migrations for multi-terabyte databases. Combines rigorous production engineering with an active R&D focus on low-latency systems, eBPF networking, and autonomous data pipelines.
AUTODOC
Ukraine (Remote)
Progressed from DevOps Engineer to Senior DevOps Engineer; company scaled from ~€500M to €1.8B+ in annual revenue during tenure
Incident Response & Reliability
Traced a Kubernetes API server dropping 82% of requests to a misconfigured admission-policy controller generating 14,000 malformed calls/second — not the initially suspected node upgrade.
Diagnosed a production database outage to long-running analytical queries exhausting replica CPU; executed emergency vertical scaling while shipping the permanent fix.
Remediated two consecutive CVSS 10.0 RCE vulnerabilities via GitOps-driven image updates, with zero service disruption.
Led the observability migration from New Relic to Grafana Cloud via OpenTelemetry, evaluating 20+ vendor criteria to remove collector-level vendor lock-in across production workloads.
Zero-Downtime Data Migrations (solo, live production)
Migrated 15-20TB+ of production data with zero downtime across six datastore types — including 10+ sharded MongoDB clusters and a 50+ node Elasticsearch cluster — using a repeatable pattern: auth introduced behind a proxy in optional mode, verified via live traces, then enforced with zero required application redeploys.
Migrated ClickHouse's coordination layer twice (ZooKeeper → ClickHouse Keeper → Raft) across sharded StatefulSets totaling 5,000+ GB.
Migrated 100+ MySQL replicas and 50+ CloudSQL instances bidirectionally via Cloud SQL migration jobs, alongside full Kafka cluster migrations with topic/partition state reconciliation.
Migrated 15-20+ Kafka, RabbitMQ, Gearman's,
CI/CD & GitOps at Scale
Configured and optimized 1,000+ CI/CD pipelines across GitLab CI and ~10 Jenkins instances; migrated GitLab itself into Kubernetes.
Authored 100+ Helm charts and custom shell-operators; built the Prometheus/VictoriaMetrics alerting stack from scratch.
Migrated hundreds of FluxCD GitOps setups across clusters — fully zero-downtime, orchestrated via Terraform/Terragrunt.
API Gateways & Traffic Management
Eliminated 82% of API request drops and resolved critical performance degradation by diagnosing and fixing a misconfigured admission-policy controller generating 14,000 malformed calls/second on Kubernetes (GKE).
Achieved zero downtime during live production migrations of 15–20TB+ of data across 6 datastore types (including 10+ sharded MongoDB clusters and a 50+ node Elasticsearch cluster) by implementing proxy-based authentication layers with zero app redeploys.
Eliminated coordination bottlenecks by successfully migrating ClickHouse's coordination layer (ZooKeeper → ClickHouse Keeper → Raft) across sharded StatefulSets totaling 5,000+ GB.
Optimized global inbound traffic and enforced rate-limiting by configuring and scaling advanced API gateways (Kong, Traefik, Envoy/Istio, and Apache APISIX) across production microservices.
Maintained zero-disruption security posture by rapidly remediating two consecutive CVSS 10.0 RCE vulnerabilities using automated GitOps (FluxCD) image updates.
Successfully migrated 100+ MySQL replicas and 50+ CloudSQL instances bidirectionally, alongside full Kafka cluster migrations with complete topic/partition state reconciliation.
Accelerated deployment cycles by configuring and optimizing 1,000+ CI/CD pipelines across GitLab CI and Jenkins, and migrating GitLab core directly into Kubernetes.
Unified global observability by migrating infrastructure monitoring from New Relic to Grafana Cloud via OpenTelemetry (OTel), removing vendor lock-in across production workloads.
Bachelor of Applied Mathematics / Bachelor of Theology, Philosophy & Religious Studies
Professional Certifications (In Progress)
Built a provider-agnostic infrastructure platform using Crossplane Compositions/XRDs and Cluster API (CAPI) + Talos OS for immutable "Cluster-as-Workload" lifecycle management across bare-metal, AWS, and Hetzner.
Secured dozens of distributed clusters with no public cluster IPs by routing all ingress through Cloudflare Tunnels and enforcing continuous state reconciliation via Crossplane.
Achieved 1–3ms end-to-end state resolution using in-memory state synchronizers, Multicall3 chunking, and weighted directed graph search algorithms (DFS/Bellman-Ford) mapped via constant-product math models.
Colocated sequencer execution on AWS Ohio using eBPF / AF_XDP kernel-bypass networking and high-IOPS dual-NVMe RAID-0 storage, paired with gas-optimized Solidity flash-loan contracts (Balancer / Aave / Uniswap V3).
Deployed 20+ production-grade Cloudflare Workers utilizing D1, KV, Vectorize, and Workers AI to ingest career data, strip PII/NDA content, and serve real-time Q&A for technical evaluators.
fully automating AWS Auto Scaling Groups (ASG) via Amazon EventBridge event gateways to manage resilient Spot instance lifecycles.
dynamically intercepts 2-minute EC2 Spot Instance Interruption Warnings and Rebalance Recommendations, leveraging AWS Spot Instance Advisor interruption rates (popularity/stability scores) and capacity-optimized allocation strategies to automatically select cheaper yet highly stable instance pools.
trigger graceful shutdowns upon termination signals, securely dumping critical node/bot state snapshots through Cloudflare tunnels to edge storage for instant recovery on newly provisioned replacement instances