Senior Network Engineer
Job description
About the role
Together AI is seeking a Senior Network Engineer to build and manage the global infrastructure powering our AI compute platforms. You will handle the full lifecycle of high-performance data center networks while collaborating across Linux, Kubernetes, and application environments. In this capacity, you will own the design, implementation, and optimization of the network fabrics that directly drive model training and inference performance. You will work closely with infrastructure and platform teams to ensure that network solutions align with demanding AI and machine learning workloads. The role requires deep expertise in both traditional and modern data center networking technologies. You will be responsible for translating complex business requirements into robust, scalable network architectures. Ultimately, your work will ensure that our infrastructure remains resilient, performant, and capable of supporting rapid innovation.
Key facts
What you'll do
- Architect and maintain global, multi-vendor network topologies that deliver the high bandwidth and low latency required for AI compute workloads.
- Investigate and resolve intricate connectivity and performance anomalies using logs, telemetry data, and packet captures to minimize service impact.
- Engineer automation frameworks and operational tooling to increase system reliability while reducing repetitive manual intervention.
- Conduct rigorous design reviews to verify that network solutions adhere to strict security, scalability, and availability benchmarks.
- Assess and recommend new hardware, optical components, and software features for evaluation and safe production deployment.
- Codify and maintain standard operating procedures for incident response, change management, and comprehensive monitoring strategies.
- Partner with systems and security engineers to troubleshoot and resolve cross-functional technical challenges that span multiple layers of the stack.
- Collaborate with platform teams to optimize network configurations for Kubernetes container orchestration and microservices communication.
- Guide the selection and integration of emerging networking technologies that support next-generation AI and HPC requirements.
- Serve as a technical authority during on-call rotations, leading escalations and directing remediation efforts for critical network outages.
Requirements
- Bring a minimum of 8 years of hands-on experience designing, operating, or troubleshooting large-scale data center, cloud, service-provider, or high-performance computing networks.
- Demonstrate mastery of core networking protocols including TCP/IP, BGP, OSPF, VXLAN, EVPN, ECMP, and QoS in complex operational environments.
- Show proven ability to design and manage multi-tenant network environments using VRFs, VLANs, network overlays, and security segmentation strategies.
- Exhibit hands-on familiarity with industry-leading network hardware from Arista, Cisco, Juniper, and NVIDIA within production settings.
- Display strong proficiency in Linux administration and fluency with command-line troubleshooting utilities such as tcpdump, Wireshark, MTR, curl, and nmap.
- Demonstrate the capability to manage network automation using Python or Ansible while adhering to Git-based version control workflows.
- Possess practical experience with Kubernetes networking concepts, including Container Network Interfaces, pod networking, and Kubernetes services.
- Show foundational knowledge of high-speed data transport technologies such as RDMA, RoCE, or InfiniBand and their implications for network design.
- Bring experience designing or operating network infrastructure for cloud environments in AWS, GCP, or Microsoft Azure.
Nice to have
- Bring direct operational experience with RoCE or InfiniBand fabrics in live environments.
- Demonstrate a background in supporting GPU-intensive clusters, high-performance computing platforms, or distributed storage systems.
- Show deep understanding of the specific traffic patterns and performance characteristics associated with AI training and inference pipelines.
- Have a track record of managing network infrastructure at scale, spanning thousands of devices distributed across multiple global regions.
- Show familiarity with techniques for validating infrastructure code and automated solutions generated by AI tools.
Skills & tools
- Python, Ansible, Git, Linux, Kubernetes, AWS, GCP, Azure, BGP, OSPF, VXLAN, EVPN, ECMP, QoS, Arista, Cisco, Juniper, NVIDIA, Wireshark, tcpdump, RoCE, InfiniBand.
Practical notes
- This is a full-time position based in San Francisco.
- Compensation components include base salary, equity awards, and comprehensive benefits.
- Salary levels are determined by considering individual experience, demonstrated skills, and geographic location.
- Together AI maintains a policy of equal opportunity for all employees and applicants.