Senior DevOps Engineer
Job description
About the role
Alpaca is on the lookout for an experienced DevOps Engineer to join our team and help manage the infrastructure that supports our global trading operations. In this pivotal position, you will play a key role in ensuring the scalability, reliability, and overall performance of our essential financial services. This opportunity allows for a high degree of independence in crafting solutions and directly influences the evolution of our platform.
Key facts
What you'll do
- Design and enhance our infrastructure on Google Cloud Platform (GCP), focusing on networking, security, and high-availability configurations, all managed through Terraform in accordance with GitOps methodologies.
- Create and oversee CI/CD pipelines for infrastructure modifications, implementing policy-as-code, drift detection, and secure deployment techniques.
- Improve our Platform-as-a-Product offerings by developing self-service tools and optimizing processes for engineering teams.
- Advance our observability capabilities, including metrics, logs, traces, and alerts, utilizing tools such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
- Manage and support Google Kubernetes Engine (GKE) clusters and related services, including message brokers like RabbitMQ and IBM MQ, as well as various data storage solutions.
- Engage in a Follow-The-Sun on-call rotation, responding to alerts and incidents while facilitating blameless post-mortems to enhance our processes.
- Incorporate Site Reliability Engineering (SRE) practices, such as defining Service Level Indicators/Objectives (SLIs/SLOs) and conducting capacity planning, into both infrastructure development and operational activities.
Requirements
- At least 5 years of experience in DevOps, Platform Engineering, or SRE roles, with a solid track record of managing extensive, highly available systems.
- In-depth hands-on experience in designing cloud infrastructure on Google Cloud Platform (GCP), including aspects like landing zones, networking, and high-availability architectures.
- Expertise in Infrastructure-as-Code with Terraform, including structuring large codebases and applying GitOps and least-privilege principles effectively.
- Proven experience in building CI/CD pipelines for infrastructure-as-code, focusing on automated planning, code reviews, policy enforcement, and secure rollouts.
- Significant production experience with Kubernetes, ideally GKE, and deploying applications using Helm.
- Strong understanding of cloud networking principles (VPCs, routing, load balancing, DNS, TLS) and the ability to troubleshoot connectivity challenges.
- Practical experience with observability tools such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager for managing metrics, logs, traces, and alerts.
- Familiarity with data storage systems like PostgreSQL and message brokers (e.g., RabbitMQ), including operational and troubleshooting skills in production environments.
- A solid understanding of SRE principles, including SLOs, error budgets, and capacity planning, along with a Platform-as-a-Product approach.
- Extensive experience in incident management, covering everything from declaration and debugging to escalation, documentation, and driving improvements post-incident.
- Willingness to participate in an APAC-based Follow-The-Sun on-call schedule and collaborate effectively in a distributed, asynchronous team setting.
Nice to have
- Experience with policy-as-code tools such as OPA or Conftest, and IaC quality tools like Checkov or tflint.
- Familiarity with managing Terraform state, module registries, and versioning at scale.
- Experience in building self-service developer platforms or internal golden paths using tools like Backstage.
- Exposure to incident management tools like Rootly or the Alloy collector.
- Working knowledge of Go for automation and tooling development.
- Strong foundational knowledge in Linux (Debian/Ubuntu) and container technologies (Docker/containerd).
- Experience with security and compliance in regulated environments, including secrets management and audit logging.
- Familiarity with the financial trading, brokerage, or fintech sectors, especially in relation to low-latency systems.
Skills & tools
- Google Cloud Platform (GCP)
- Terraform
- GitOps
- Kubernetes (GKE)
- Helm
- Prometheus
- Thanos
- Grafana
- Loki
- Tempo
- Alertmanager
- PostgreSQL
- RabbitMQ
- IBM MQ
Practical notes
- Competitive salary along with stock options.
- Comprehensive health benefits.
- One-time home-office setup allowance of USD $500.
- Monthly stipend of USD $150 provided via Brex Card.
About the company
A global farm business cares for domesticated camelids from South America. Origin herds once lived on high ground across Peru, Bolivia, Ecuador, and Chile. Today ranches in many regions welcome thousands of new births each year. Strong interest exists in North America, Europe, and Australia. Gentle animals form the focus of daily work. Stable routines guide animal health and development. Long standing practices shape modern standards. Knowledge sharing supports communities everywhere.