Infrastructure Engineer
Job description
About the role
You will own the design, operation, and evolution of the storage infrastructure that powers large-scale AI and HPC workloads at Lightning AI. You will partner closely with hardware, platform, and application teams to ensure the data plane meets extreme demands for throughput, latency, and reliability. You will drive automation and observability across our storage stack, from low-level device and network tuning to high-level service orchestration. You will implement resilient architectures that balance performance, durability, and cost across heterogeneous storage media. You will own runbooks, incident response, and capacity planning to keep critical training and inference pipelines online. You will contribute to open source projects where applicable and translate production constraints into clear requirements for vendors and partners. You will mentor junior engineers and elevate engineering standards through code reviews, documentation, and design discussions.
Key facts
What you'll do
- Operate and scale distributed storage systems, including VAST and S3-compatible object storage such as Ceph.
- Automate provisioning, upgrades, and recovery workflows for storage clusters to reduce manual toil and increase reliability.
- Tune hardware, network, and kernel parameters to maximize throughput and minimize latency for data-intensive AI workloads.
- Own monitoring, alerting, and troubleshooting across the storage stack to ensure high availability and rapid incident response.
- Analyze performance metrics and workload patterns to guide capacity planning and infrastructure investments.
- Collaborate with cross-functional teams to translate application requirements into storage architectures and service-level objectives.
- Implement and maintain backup, retention, and data integrity mechanisms aligned with security and compliance needs.
- Evaluate and integrate new storage technologies and vendors to improve cost-efficiency and capability.
- Document system designs, operational procedures, and failure modes to support on-call and handoffs.
- Participate in on-call rotations to support production storage services across global data centers.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- 5+ years of experience operating storage systems or infrastructure at scale.
- Deep understanding of storage architectures, including object, block, and file systems.
- Hands-on experience with distributed storage systems such as Ceph, S3-compatible services, or similar.
- Experience with hardware components including SSDs, NVMe, and networking devices in data center environments.
- Strong scripting and programming skills in Python or Go for automation and instrumentation.
- Proven ability to troubleshoot complex performance, reliability, and data integrity issues.
- Comfort working in fast-paced, ambiguous environments where priorities shift quickly.
Nice to have
- Experience with open source storage projects or contributing upstream.
- Background in high-performance computing or AI/ML infrastructure workloads.
- Familiarity with cloud storage services and hybrid cloud patterns.
Practical notes
- This role requires a minimum of 2 in-office days per week and occasional team and company offsites.
- We are not able to provide visa sponsorship for this position at this time.
- Compensation details are not specified in the source information provided.
- The engagement type and exact hours are to be confirmed from the source.
- Locations are limited to the offices in New York City, San Francisco, Seattle, and London.
- All communication and collaboration will align with our documented ways of working, including openness, ownership, and continuous improvement.
- Decisions on tooling, architecture, and process will be made with consideration for long-term scalability and simplicity.
- You will be expected to raise the bar over time by learning from feedback, challenging yourself, and focusing on high-impact work.
About Lightning AI
Lightning AI is hiring for Infrastructure Engineer. The listing location is London, England, United Kingdom; New York, New York, United States; San Francisco, California, United States; Seattle, Washington, United States.
This Infrastructure Engineer opening is posted for London, England, United Kingdom; New York, New York, United States; San Francisco, California, United States; Seattle, Washington, United States.