Senior Infrastructure Engineer - IDC
BinanceAsiaFull-time Onsite or Remote1w ago
KubernetesWeb3BlockchainSecurityFinanceEngineeringInfrastructureremotecurated-jd
Job description
Senior Infrastructure Engineer - IDC at Binance
About the role
Join Binance to help construct and manage the physical and software infrastructure that powers our global operations. You will be instrumental in building out our data center presence, ensuring high availability and performance for our critical systems. This role requires a deep understanding of both hardware and software infrastructure.
Key facts
What you'll do
- Plan and implement the physical arrangement of equipment within data halls, including network pathways and overall system setup.
- Develop and deploy a secure, separate management network for out-of-band access, covering BMC, IPMI, and Redfish functionalities.
- Collaborate with data center providers and hardware suppliers for server deployment, network connections, and on-site support.
- Create automated processes for the rapid deployment of operating systems onto new hardware.
- Establish and maintain production-ready, self-hosted Kubernetes environments, including core components like Cilium and GitOps tools.
- Manage the entire lifecycle of Kubernetes clusters, including updates, scaling, node maintenance, and incident resolution.
- Document operational procedures, create runbooks, and conduct post-incident reviews, while participating in a 24x7 on-call schedule.
- Set up connectivity between our data centers and public cloud environments using VPNs, direct connections, or peering.
- Prepare our infrastructure to support large-scale GPU and VM workloads alongside containerized applications.
Requirements
- A minimum of 5 years of experience in data center, infrastructure, or platform engineering.
- Proven ability to build physical data center infrastructure from the ground up, including rack design, network configuration, server setup, and vendor coordination.
- Practical experience in designing and managing out-of-band management networks (BMC, IPMI, Redfish).
- Demonstrated experience in deploying and operating self-hosted Kubernetes clusters in production, with the ability to troubleshoot cluster-level issues.
- Strong proficiency in Linux system administration and performance optimization, covering kernel, networking, and storage.
- Experience with bare-metal automation tools such as MAAS, Tinkerbell, or Cluster API.
- Expertise in infrastructure as code tools like Terraform and Ansible, along with proficiency in at least one scripting language (Python, Go, or Bash).
- Familiarity with Cisco networking concepts and configurations, including VLAN, LACP/LAG, BGP, and ACLs.
- Experience configuring Palo Alto firewalls.
- Experience with enterprise storage solutions from vendors like NetApp, Dell EMC, or Pure Storage.
- Fluency in either English or Mandarin.
Nice to have
- Experience with high-density racks (30kW+) and high-speed networking (400G+).
- Familiarity with immutable operating systems like Talos Linux, Flatcar, or Bottlerocket.
- Proficiency in both AWS and GCP, including experience with cross-cloud data migration.
- Experience building storage clusters or big data/offline data processing clusters.
- Experience with virtualization technologies such as KubeVirt.
- Exposure to NVIDIA GPU Operator and Kubernetes GPU workloads.
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Security Specialist (CKS).
- Cisco Certified Network Professional (CCNP) or Cisco Certified Internetwork Expert (CCIE) certification.
- Professional-level Japanese language skills.
Skills & tools
- Linux Systems Administration
- Kubernetes
- Terraform
- Ansible
- Python / Go / Bash
- Cisco Networking (VLAN, LACP/LAG, BGP, ACL)
- Palo Alto Firewalls
- MAAS / Tinkerbell / Cluster API
- Cilium
- ArgoCD
- NetApp / Dell EMC / Pure Storage
- BMC / IPMI / Redfish
- KubeVirt
- NVIDIA GPU Operator
Practical notes
Binance offers competitive compensation and benefits. This role may involve participation in a 24x7 on-call rotation.