AI infrastructure System Engineer Bangalore
Job description
AI Infrastructure Systems Engineer at Together AI.
About the role
This role designs and operates infrastructure for frontier model training and inference. The work emphasizes software driven operations instead of manual dashboards.
Software engineers turn product ideas into working code. Engineers work in small teams, review each other's work, and ship in small batches. Most teams follow agile practices such as sprints and daily standups. Engineers also write tests, fix bugs, and improve performance. The field values clear communication as much as technical skill. Engineers spend part of every week on planning, code review, and debugging, not just writing new code. The ability to explain a technical decision in plain words separates strong engineers from the rest.
Key facts
What you'll do
Automated systems handle provisioning, validation, deployment, upgrading, repair, and retirement of GPU clusters with minimal human intervention.
Internal platforms and developer tools shift infrastructure management from manual operations to software-defined programmable logic.
Automation improves deployment velocity, reliability, and operational efficiency to shorten lead times and reduce operational toil.
Requirements
3+ years building distributed systems, infrastructure platforms, or large scale backend software with experience in complex systems and production operations.
Strong software engineering skills in Python, Go, or Rust with an expectation of code quality and maintainability in production systems.
Experience building platforms, automation systems, or developer infrastructure that reduces manual effort.
Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies to manage environments and workflows at scale.
Understanding problems across hardware and software to connect low level signals to high level outcomes through analysis that spans layers and components.
A passion for solving complex infrastructure challenges through software, favoring robustness, scale, and simplicity.
An automation-first mindset where repeated tasks trigger building a system to eliminate them, targeting error prone and tedious steps.
Nice to have
GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch experience.
InfiniBand or RoCE networking knowledge.
Bare-metal provisioning and lifecycle management background.
Large-scale AI training or inference cluster operations.
Hardware health monitoring and predictive failure detection work.
Distributed storage systems experience.
AI agents and autonomous infrastructure operations familiarity.
Practical notes
This role is based in Bangalore, India. Equal opportunity applies to all candidates. Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
The role centers on software defined infrastructure for AI compute. Work directly impacts large scale GPU fleet performance. The focus is on autonomous operations and system wide reliability. Engineers build platforms that replace repetitive manual tasks. The environment values deep technical problem solving at scale.
Questions to ask
Good questions to ask the employer in the interview: what does success look like in the first six months, how is the team structured, what is the current biggest challenge, and how are decisions made. Asking about growth paths and the review process is also well received. Employers expect questions, and good ones show preparation.
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels. Some engineers move into management and lead teams of five to twenty people. Others stay on the technical track. Growth follows demonstrated impact, not tenure alone. A typical engineering ladder has clear levels with defined expectations for scope, quality, and mentorship. Moving up usually requires owning outcomes end to end rather than completing assigned tickets.