Site Reliability Engineer
Job description
Site Reliability Engineer at PhonePe.
About the role
PhonePe is looking for a System Engineer to manage the infrastructure of our large-scale fintech platform. You will focus on maintaining high availability and performance across our services while working with a variety of open source and cloud technologies.
Key facts
What you'll do
- Maintain service uptime and availability through established IT operations and best practices.
- Participate in the team on-call rotation.
- Manage Linux/Unix environments and private cloud setups.
- Automate tasks and handle configuration management.
- Collaborate on infrastructure projects to support our payment and financial service ecosystem.
Requirements
- Minimum 4 years of professional experience in Linux/Unix administration.
- Proficiency in computer networking, specifically with IP, iptables, and IPsec.
- Hands-on experience with MySQL databases.
- Scripting and coding ability in Python, Perl, or Golang.
- Experience with SaltStack for automation.
- Ability to communicate clearly in English.
Nice to have
- Experience with KVM/QEMU for cloud services on Linux.
- Knowledge of container orchestration.
- Familiarity with NoSQL databases, specifically Aerospike.
- Experience with Galera clusters.
- Proficiency in Perl or Golang.
- Practical experience with data center operations and MariaDB or Percona.
Skills & tools
- Linux/Unix
- Networking (IP, iptables, IPsec)
- MySQL
- Python, Perl, Golang
- SaltStack
- Private cloud environments
Practical notes
Benefits include medical, critical illness, accidental, and life insurance. We provide wellness programs, parental support (maternity, paternity, adoption, and day-care), and retirement benefits like PF, Gratuity, and NPS. Additional perks include higher education assistance, car leases, and relocation support. PhonePe is an equal opportunity employer; candidates requiring disability-related accommodations during the hiring process should use the provided form link.