Infrastructure Engineer
Job description
com.
About the role
You will own the reliability and scalability of our core Kubernetes infrastructure that underpins all Kraken trading and business operations. You will collaborate daily with SRE, Networking, and Security teams to design, operate, and evolve the platform that runs our most critical workloads. This role provides hands-on experience across cluster operations, distributed systems debugging, and infrastructure automation that directly impacts global financial services. You will triage complex infrastructure issues, implement automation to reduce manual toil, and participate in on-call rotations to ensure platform stability. You will help build and operate Kubernetes clusters across multiple datacenters while contributing to a platform that advances open finance. You will leverage AI tools and scripting skills to solve operational problems and improve efficiency for the broader engineering organization.
Key facts
What you'll do
- Provide technical guidance to stakeholders with excellent communication, removing blockers for engineering teams using Kubernetes.
- Build and operate Kubernetes clusters across multiple datacenters, ensuring high availability and performance.
- Triage and organize infrastructure issues, learning our systems, best practices, and operational runbooks in depth.
- Work with senior engineers to solve complex problems, reducing operational burden from the broader platform and product teams.
- Build automation and tooling using scripting or programming languages to improve team efficiency and reduce repetitive manual tasks.
- Participate in on-call rotation to maintain platform reliability and respond to incidents promptly and effectively.
- Gain hands-on experience across Kubernetes cluster operations, deployments, networking, and scalability in a production environment.
- Leverage AI tools and agents such as Claude and OpenAI to efficiently deliver business value and accelerate issue resolution.
- Apply strong Linux systems knowledge at the shell, including processes, networking basics, file systems, and permissions.
- Demonstrate understanding of networking fundamentals such as TCP/IP, DNS, ports, IP addressing, basic routing, TLS, and PKI concepts.
- Contribute to the evolution of our Kubernetes platform as a core component of our next-generation infrastructure strategy.
- Utilize familiarity with orchestration and advanced Kubernetes networking, including Cilium, to optimize platform capabilities.
- Apply understanding of distributed systems fundamentals to design and troubleshoot scalable infrastructure solutions.
- Use scripting or programming experience in Python, Bash, Go, or similar languages to automate infrastructure tasks and improve observability.
Requirements
- Bring 3+ years of proven experience as a Site Reliability Engineer, Infrastructure/Platform/DevOps Engineer, Software Engineer, or similar roles.
- Show strong passion for providing technical guidance to stakeholders and maintaining a customer-focused approach to resolving Kubernetes-related requests.
- Demonstrate a strong understanding of distributed systems fundamentals and how they apply to real-world platform challenges.
- Exhibit strong Linux systems knowledge, including comfort at the shell, understanding of processes, networking basics, file systems, and permissions.
- Possess networking fundamentals such as TCP/IP, DNS, ports, IP addressing, basic routing, TLS, and PKI concepts.
- Have a solid understanding of Kubernetes internals, including the control plane, scheduling, networking, and storage concepts.
- Display familiarity with orchestration or advanced Kubernetes networking solutions such as Cilium.
- Show ability to leverage AI tools and agents like Claude and OpenAI to efficiently deliver business value and improve workflows.
- Have scripting or programming experience in languages such as Python, Bash, Go, or similar to automate tasks and solve infrastructure problems.
- Maintain reliability and professionalism while participating in on-call responsibilities and incident response.
- Communicate clearly and effectively with cross-functional teams to unblock engineering work and support platform adoption.
- Follow best practices and contribute to improvements in operational processes and documentation.
- Adhere to company policies and compliance standards relevant to financial infrastructure and open finance platforms.
- Be comfortable working in a fast-paced environment where priorities can shift based on business and platform needs.
- Collaborate closely with senior engineers to grow skills and advance platform reliability and scalability objectives.
Nice to have
- Bring working experience running Kubernetes clusters in production environments.
- Demonstrate exposure to cloud platforms such as AWS, GCP, or Azure, or to on-premises infrastructure.
- Show knowledge of Terraform or other Infrastructure as Code tools to automate environment provisioning.
- Have experience with storage systems such as Ceph or Rook for persistent data management.
Practical notes
This role is based in LATAM and is offered as a full-time engagement. No specific travel or visa requirements are outlined in the source material, and no application deadline is stated, so applications are accepted on an ongoing basis.