Network Operations Center
Job description
About the role
Lightning AI is hiring Network Operations Center Analysts to maintain 24/7 oversight of our high-performance computing data centers. You will serve as the primary technical contact for monitoring infrastructure, diagnosing system health, and resolving issues across our compute and network environments. In this capacity, you will act as the central point for detecting, analyzing, and coordinating responses to infrastructure events. The role requires a proactive mindset focused on maintaining the stability and performance of critical systems around the clock. You will work within a structured shift schedule to ensure continuous coverage and rapid response times. Success in this position depends on your ability to manage routine monitoring tasks and handle unexpected disruptions with clarity. Your work will directly impact the reliability and availability of the platforms supporting our compute and network operations.
Key facts
What you'll do
- Monitor telemetry data, dashboards, and alerts to identify system anomalies across compute and network environments.
- Conduct independent technical diagnostics on Linux systems, hardware health, and network connectivity using command-line tools and centralized logs.
- Troubleshoot network-layer problems such as routing anomalies, interface errors, and end-to-end connectivity failures.
- Triage and escalate incidents to specialized teams including SRE, hardware, and network groups with comprehensive documentation.
- Maintain accurate incident tickets that capture diagnostic steps, observations, and chronological event details for audit and review.
- Analyze alert patterns and historical data to suggest improvements for monitoring reliability, coverage, and automation opportunities.
- Collaborate with engineering teams to validate configuration changes and verify remediation actions through observation.
- Perform routine verification checks on infrastructure components to ensure systems are operating within defined thresholds.
- Document standard procedures and runbooks to support consistent responses to recurring issues and on-call requirements.
- Contribute to continuous improvement initiatives by sharing insights from incident reviews and operational performance metrics.
Requirements
- Proficiency in Linux command-line operations and system log analysis using standard tools and utilities.
- Working knowledge of networking fundamentals including TCP/IP, DNS, routing protocols, and the use of diagnostic tools such as tcpdump, netstat, ping, and traceroute.
- Capacity to diagnose complex technical issues independently and make reasoned decisions in ambiguous or high-pressure scenarios.
- Strong technical writing and communication skills for creating clear, concise, and accurate incident documentation and status updates.
- Ability to work a 24/7 schedule, including overnight shifts, weekends, and rotating coverage across multiple time zones as required.
- Willingness to follow established on-call procedures and respond promptly to escalations and operational alerts.
- Adherence to organizational standards for security, data handling, and operational procedures in a regulated environment.
- Physical capability to perform onsite work at data center facilities, including extended periods of sitting or standing while working at consoles.
Nice to have
- Experience with monitoring and visualization platforms such as Datadog, Prometheus, or Grafana for operational analysis.
- Background working with GPU-based or HPC infrastructure, including job scheduling systems and high-throughput storage architectures.
- Scripting abilities in Python or Bash to automate repetitive tasks, parse logs, and build custom diagnostic workflows.
Practical notes
- This position is based onsite at our data centers in Fort Worth, Texas, Lisle, Illinois, and Quincy, Washington.
- We cannot provide visa sponsorship for this role.
- Compensation includes a base salary ranging from $85,000 to $100,000 USD, discretionary bonus, equity in the form of RSUs, and a comprehensive benefits package.
- Benefits coverage includes medical, dental, and vision insurance, along with 401(k) matching contributions.
- Additional benefits consist of unlimited PTO, a structured winter break period, parental leave policies, and a professional development allowance for training and certifications.
- The role also features wellness stipends and a sabbatical program available after four years of continuous service.
- Applicants must be able to commit to the full duration of the engagement with consistent attendance and reliability.
- Deadlines for application review will be managed on a rolling basis until the position is filled.
- Travel is not required for this role, and all work activities are conducted at the designated on-site locations.
- Candidates must be legally authorized to work in the United States without requiring sponsorship from the employer.