Senior DevOps Engineer
AlpacaRemote (Japan - APAC)2w ago
Engineeringremotecurated-jd
Job description
Senior DevOps Engineer at Alpaca
About the role
Alpaca is seeking a seasoned DevOps Engineer to build and maintain the infrastructure powering our global trading systems. You will be instrumental in ensuring scalability, reliability, and performance for critical financial services. This role offers significant autonomy in designing solutions and a direct impact on shaping our platform's future.
Key facts
What you'll do
- Architect and evolve our Google Cloud Platform (GCP) infrastructure, including networking, security, and high-availability setups, managed entirely through Terraform following GitOps principles.
- Develop and manage CI/CD pipelines for infrastructure changes, incorporating policy-as-code, drift detection, and safe deployment strategies.
- Enhance our Platform-as-a-Product offerings, creating self-service tools and streamlined processes for engineering teams.
- Strengthen our observability suite, encompassing metrics, logs, traces, and alerts using tools like Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
- Operate and maintain Google Kubernetes Engine (GKE) clusters and associated services, including message brokers like RabbitMQ and IBM MQ, and data stores.
- Participate in a Follow-The-Sun on-call rotation, handling alerts, incidents, and driving blameless post-mortems.
- Integrate Site Reliability Engineering (SRE) practices, such as Service Level Indicators/Objectives (SLIs/SLOs) and capacity planning, into infrastructure development and operations.
Requirements
- A minimum of 5 years of experience in DevOps, Platform Engineering, or SRE roles, with a proven history of managing large-scale, highly available systems.
- Extensive hands-on experience designing cloud infrastructure on Google Cloud Platform (GCP), covering landing zones, networking, and high-availability topologies.
- Proficient in Infrastructure-as-Code using Terraform, with experience structuring large codebases and implementing GitOps and least-privilege principles.
- Demonstrated experience building CI/CD pipelines for infrastructure-as-code, including automated planning, code review, policy enforcement, and safe rollouts.
- Significant production experience with Kubernetes, preferably GKE, and deploying applications using Helm.
- Solid understanding of cloud networking fundamentals (VPCs, routing, load balancing, DNS, TLS) and troubleshooting connectivity issues.
- Practical experience with observability tools such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager for metrics, logs, traces, and alerts.
- Operator-level familiarity with data stores like PostgreSQL and message brokers (e.g., RabbitMQ), including their operation and troubleshooting in production.
- A good grasp of SRE principles, including SLOs, error budgets, and capacity planning, coupled with a Platform-as-a-Product mindset.
- Comprehensive experience with incident management, from declaration and debugging to escalation, documentation, and driving post-incident improvements.
- Ability and willingness to participate in an APAC-based Follow-The-Sun on-call schedule and collaborate effectively in a distributed, asynchronous team environment.
Nice to have
- Experience with policy-as-code tools like OPA or Conftest, and IaC quality tools such as Checkov or tflint.
- Familiarity with managing Terraform state, module registries, and versioning at scale.
- Experience building self-service developer platforms or internal golden paths using tools like Backstage.
- Exposure to the Alloy collector or incident management tools like Rootly.
- Working knowledge of Go for automation and tooling development.
- Strong fundamentals in Linux (Debian/Ubuntu) and container technologies (Docker/containerd).
- Experience with security and compliance in regulated environments, including secrets management and audit logging.
- Familiarity with financial trading, brokerage, or fintech domains, and low-latency systems.
Skills & tools
- GCP
- Terraform
- GitOps
- Kubernetes (GKE)
- Helm
- Prometheus
- Thanos
- Grafana
- Loki
- Tempo
- Alertmanager
- PostgreSQL
- RabbitMQ
- IBM MQ
Practical notes
- Competitive salary and stock options.
- Health benefits.
- USD $500 one-time home-office setup allowance.
- USD $150 monthly stipend via Brex Card.