Datacenter Infrastructure Specialist
Job description
About the role
RunPod is on the lookout for a Datacenter Infrastructure Specialist to take on a role in overseeing our extensive global fleet. This position involves reporting directly to the Manager of Infrastructure Capacity & Management, serving as a vital technical liaison between our hardware partners and internal engineering teams. Your specialized knowledge will be instrumental in maintaining the operational efficiency of our growing GPU fleet, ensuring it meets the rigorous demands of AI workloads.
Key facts
What you'll do
- Assess and benchmark new hardware to confirm that deployments meet RunPod's specifications for distributed AI and machine learning workloads.
- Continuously monitor fleet performance to identify potential issues, auditing downtime and supplying essential technical data to maintain customer service level agreements.
- Implement an AI-centric approach to optimize operations, utilizing large language models (LLMs) and AI agents to automate network troubleshooting and develop dynamic runbooks.
- Oversee technical incident communications during outages, ensuring clear updates and actionable solutions are provided.
- Assist in the development of our infrastructure partners by offering technical support and guidance.
- Collaborate with internal teams to enhance system performance and reliability, ensuring alignment with operational goals.
- Participate in post-incident reviews to analyze performance and identify areas for improvement.
- Stay updated on industry trends and emerging technologies to recommend enhancements to our infrastructure.
Requirements
- Professional Background: 3 to 5 years of experience in infrastructure operations, systems reliability, or datacenter engineering is essential.
- Datacenter Networking: In-depth understanding of standard datacenter networking and troubleshooting performance issues, with a preference for experience with RDMA, InfiniBand, or RoCE.
- GPU & AI Stack: Hands-on experience with the NVIDIA Software Stack, including driver installation and performance utilities, along with knowledge of multi-node performance tuning.
- Systems & Diagnostics: Proficient in Linux system administration and containerization (Docker), capable of performing system-level troubleshooting and performance optimization.
- Effective Communication: Strong written and verbal communication skills to articulate hardware or networking issues to both technical partners and internal teams.
- Operational Flexibility: Willingness to engage in an on-call rotation as our global fleet expands.
- Strategic Problem-Solver: Detail-oriented and proactive in identifying potential issues before they impact customers.
Nice to have
- Startup Experience: Familiarity with dynamic environments and contributing to the development of operational workflows.
- HPC Exposure: Experience in managing or optimizing bare-metal High-Performance Computing environments at scale is a plus.
- Observability Tools: Knowledge of monitoring tools such as Grafana, Prometheus, or Datadog for system health oversight.
- Automation: Proficiency in programming languages like Python, Go (Golang), or Bash for automating infrastructure tasks and interfacing with internal APIs.
- A competitive base salary ranging from $120,000.00 to $160,000.00, determined by your experience, qualifications, and location during the interview process.
- Equity options in a rapidly expanding AI infrastructure company, allowing you to share in the company's success.
- Comprehensive medical, dental, and vision insurance plans, fully covering employees and partially covering dependents.
- Flexible paid time off to ensure you can recharge as needed.
- Primarily remote work with a collaborative team environment, utilizing Slack for internal communication.
- A supportive culture focused on learning and ownership, essential for scaling our operations.
- A $1,200 stipend for home office equipment to help you establish an effective workspace from day one.
RunPod is committed to creating an inclusive workplace that values diversity in all its forms. We are an equal opportunity employer and evaluate applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome all qualified candidates eligible to work in the United States; however, we are currently unable to sponsor employment visas.