Senior Software Engineer
Job description
About the role
Together AI is building an AI Acceleration Cloud designed to support the entire generative AI lifecycle from research to deployment. In this role, you will own the design and execution of critical infrastructure components that virtualize advanced ML hardware at global scale. You will partner closely with hardware, platform, and product teams to turn complex data center requirements into reliable software abstractions. Your work will directly influence the performance and reliability of both internal products and external customer workloads. You will investigate and prototype solutions for decentralized AI compute patterns and their integration into a unified control plane. You will contribute code back to the open-source Together AI platform that powers our commercial cloud. You will translate ambiguous operational challenges into durable software designs with clear interfaces and test coverage.
Key facts
What you'll do
- Develop and maintain secure backend services and low-level operators that automate data center hardware lifecycle management, including Infiniband subnet partitioning and dynamic storage provisioning policies.
- Construct the IaaS software layer for next-generation GB200 data centers, managing thousands of GPUs and their associated networking fabric with strict isolation and performance guarantees.
- Build a globally distributed, multi-exabyte object storage layer to serve massive pretraining datasets, optimizing for throughput, consistency, and cost efficiency.
- Create end-to-end observability stacks and node lifecycle management tools to ensure fault-tolerant distributed training jobs and rapid failure recovery.
- Conduct deep architectural research into decentralized AI workloads, exploring execution models, fault domains, and resilience patterns for geographically dispersed clusters.
- Contribute actively to the core open-source Together AI platform, ensuring that internal innovations are generalized and released for community use.
- Author precise developer documentation and comprehensive testing frameworks that validate system reliability under failure and scale conditions.
- Collaborate with hardware and site reliability engineers to translate data center constraints into software requirements and configuration standards.
- Design control plane components that abstract heterogeneous compute resources into uniform service endpoints for downstream orchestration.
- Implement automation for secure credential rotation, network policy enforcement, and audit logging across multi-tenant deployments.
- Optimize data plane pipelines to reduce training job startup times and improve aggregate throughput across global regions.
- Evaluate and integrate emerging standards in hardware virtualization and workload scheduling to maintain competitive infrastructure advantages.
- Partner with product teams to align infrastructure roadmaps with customer needs and emerging AI model architectures.
- Troubleshoot complex performance regressions by correlating metrics, traces, and logs across distributed infrastructure layers.
Requirements
- Hold a Bachelor's degree or equivalent practical experience in Computer Science, Engineering, or a related technical field.
- Bring 5+ years of professional software engineering experience focused on infrastructure, distributed systems, or platform engineering.
- Demonstrate 5+ years of experience producing high-performance, thoroughly tested, production-grade code that operates at scale.
- Show proven proficiency in at least one backend programming language, with a demonstrated preference for Golang in prior roles.
- Provide a track record of building and managing globally distributed micro-service architectures deployed on major cloud providers such as AWS, Azure, or GCP.
- Exhibit strong ability to produce clear design documentation and communicate effectively with both technical and non-technical stakeholders.
- Maintain solid systems knowledge regarding memory management, concurrency models, performant I/O patterns, and scaling strategies for stateless and stateful services.
- Have hands-on experience with infrastructure automation tools like Terraform or Ansible, observability stacks such as Prometheus and Grafana, and CI/CD pipelines including GitHub Actions or ArgoCD.
- Display familiarity with security best practices for multi-tenant environments, including identity management, network segmentation, and audit compliance.
- Show commitment to writing clean, maintainable code with strong testing discipline and version control practices.
- Demonstrate ownership of on-call responsibilities and the ability to respond to production incidents at scale with minimal disruption.
- Exhibit curiosity and disciplined learning to quickly master new hardware, networking, and virtualization technologies in fast-paced data center environments.
- Communicate clearly in written and verbal formats, ensuring technical tradeoffs are understood across diverse teams.
- Align with Together AI's mission to build open, efficient, and reliable infrastructure for the global AI ecosystem.
Nice to have
- Deep knowledge of Kubernetes internals, including custom operators, advanced scheduling strategies, and network or storage plugin development.
- Experience with hypervisors and virtualization technologies such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, or SR-IOV in production settings.
- Expertise in data center networking including VLAN, VXLAN, VPN, VPC architectures, and Open vSwitch or OVN implementations.
- Familiarity with Cluster API, high-performance computing paradigms, GPU virtualization techniques, and Infiniband management tools.
- Background in building IaaS or PaaS systems at scale, with clear examples of operational experience in large deployments.
- Experience with DPUs or SmartNICs, GPU programming models, NCCL collectives, or CUDA kernel optimization for data plane offload.
Skills & tools
Core toolchains and platforms emphasized for this role include Golang, Kubernetes, AWS/Azure/GCP, Terraform, Ansible, Prometheus, Grafana, GitHub Actions, ArgoCD, QEMU/KVM, Infiniband fabrics, and CUDA where applicable.
Practical notes
- This role offers flexibility regarding remote work, with options for hybrid collaboration from the San Francisco office.
- Together AI is an Equal Opportunity Employer committed to building a diverse and inclusive team.
Application instructions are provided in the official career portal with details on submitting your resume and relevant project artifacts. Please ensure your documentation reflects direct experience with the technologies listed in the requirements and nice to have sections. The compensation band is aligned with level titles in the San Francisco metropolitan area and includes equity grants reviewed periodically. Candidates are encouraged to highlight specific contributions to infrastructure reliability, performance, and developer experience in their submissions.