Senior Site Reliability Engineer
Job description
Senior Site Reliability Engineer at TikTok USDS Joint Venture
About the role
We are looking for a skilled Site Reliability Engineer to enhance the reliability of TikTok's video platform. This role involves ensuring the performance and stability of our video services, which cater to billions of users globally. If you enjoy tackling complex challenges and improving system reliability, we want you to join our team.
Key facts
What you'll do
- Oversee the reliability of TikTok's video services, focusing on publishing and distribution.
- Manage the lifecycle of production systems, including change management and emergency response.
- Monitor system performance and address incidents to uphold service level agreements, while reviewing and following up on production issues.
- Manage resource capacity for computing, storage, and network bandwidth to ensure system stability and optimize costs.
- Provide support during high-traffic events to maintain service quality.
- Develop tools, automations, and monitoring systems to enhance global infrastructure operations.
Requirements
- A Bachelor's degree in Computer Science or a related technical field, or equivalent practical experience.
- At least 5 years of experience in Site Reliability Engineering or software development for large-scale online services.
- Proficiency in at least one programming language such as C, C++, Java, Python, C#, or Go.
Nice to have
- In-depth knowledge of networking, operating systems, database systems, and container technologies.
- Familiarity with microservice architecture and experience troubleshooting large-scale distributed systems.
- Practical experience with open-source technologies like Linux, MySQL, MongoDB, Redis, and ELK.
- Experience with cloud services such as AWS, Google Cloud, or Azure is advantageous.
- A self-motivated approach with strong teamwork skills.
Skills & tools
- Proficiency in system monitoring and incident management tools.
- Familiarity with automation and orchestration tools.
- Understanding of cloud infrastructure and services.
Practical notes
This position is eligible for U.S. work visas. Benefits include medical, dental, and vision insurance from day one, a 401(k) plan with company matching, paid parental leave, and various other wellbeing benefits. Employees receive 10 paid holidays, 10 sick days, and 17 days of paid personal time per year, with accruals increasing over time. The salary for this role ranges from $177,688 to $341,734 annually, depending on qualifications and experience. Additional bonuses and stock options may be available.