Software Engineer, Infrastructure
Job description
Software Engineer, Infrastructure at Fal Ai.
About the role
Infrastructure Engineering Role at fal
fal operates as the generative media ecosystem enabling the next wave of artificial intelligence applications. The company constructs the underlying infrastructure, toolsets, and model accessibility required for teams to transition concepts into production environments at scale. fal provides a unified platform where high-performance inference, orchestration, and observability converge, making generative media deployment practical for developers and enterprises. This role exists within a market experiencing substantial growth, and fal is establishing the foundation upon which ambitious teams build their products.
Position Overview
This position involves hands-on engineering focused on sustaining the health and productivity of a large-scale GPU server fleet. The responsibility includes designing software and processes for provisioning, monitoring, diagnosing, and recovering thousands of servers. The role requires collaboration with partners when automated systems cannot resolve issues. The primary goal is to ensure the infrastructure remains reliable and efficient under demanding AI workloads.
Key Responsibilities
The position holder will construct and maintain a Python-based fleet tracking system. This system will oversee the complete server lifecycle, including contracting, procurement, target utilization, pricing, availability, health status, and RMA management. Server management tooling will be developed to automate provisioning procedures, execute health validations, perform GPU diagnostics, initiate recovery actions, and generate alerts.
Creation and maintenance of hardware health metrics, dashboards, and alerting channels will cover GPU errors, disk failures, network issues, and thermal conditions. The role requires leveraging artificial intelligence extensively to build tools that automate alerting and initiate recovery procedures across the entire fleet. Implementation and enforcement of OS-level security measures, including hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation, are essential.
Management and optimization of distributed and local storage systems will support model weights, checkpoints, and ephemeral scratch space. This includes NVMe arrays, NFS, parallel file systems, and object storage solutions. The role involves tuning Linux systems specifically for AI workloads, addressing kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and the GPU driver stack, including NVIDIA drivers, CUDA, and container runtimes.
Developing a comprehensive suite of automated error detection and recovery processes forms a core function. Collaboration with partners to resolve technical issues and align strategic direction is also required.
Position Requirements
Applicants must possess over three years of experience managing bare-metal and cloud-based server fleets at a scale exceeding 100 nodes. Strong software engineering capabilities in Python are mandatory; the role involves building production-grade tooling rather than writing scripts. Deep knowledge of Linux internals is required, encompassing boot processes, networking, storage systems, and performance analysis techniques.
Proficiency in configuration management and infrastructure-as-code using tools such as Ansible, Terraform, and cloud-init is necessary. Experience with storage technologies, including LVM, RAID, NVMe, NFS, Lustre, and GPFS, is required. The candidate must recognize hardware failure patterns for GPUs, disks, and networks based on diagnostic data. Experience in building internal observability tools and dashboards for infrastructure visibility is essential. Clear communication skills and the ability to influence technical decisions across diverse teams are required. The individual must be a self-starter who executes efficiently, takes ownership of tasks, and actively seeks improvements.
Preferred Qualifications
Familiarity with network configuration, diagnostics, and traffic capture analysis is advantageous. Experience with NVIDIA GPU infrastructure, including driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, and InfiniBand/RoCEv2, is considered valuable. Experience with AMD GPU infrastructure and related tooling is also beneficial.
Background in bare metal and virtualized provisioning using PXE, Kickstart, libvirt, and Qemu/KVM is preferred. Understanding of compliance frameworks relevant to cloud providers, such as SOC 2 and ISO 27001, is also a plus.
Location and Compensation
Benefits and Offerings
fal offers work that is interesting and challenging. The role provides significant opportunities for learning and professional growth. Regular team events and offsites are organized to support team cohesion.
Role Context and Data
This role replaces the previous "Software Engineer, Infrastructure" position. Key facts regarding the new position include a requirement for 3+ years managing large server fleets. The successful candidate will write scalable Python production tooling, possess deep Linux internals knowledge, and practice infrastructure as code. Proficiency in storage technologies, recognition of failure patterns, and experience creating observability tools are required. Clear communication and ownership are emphasized. Preferred skills include network diagnostics, NVIDIA and AMD GPU stack experience, bare metal and VM provisioning knowledge, and compliance practices. The role involves using Python, Linux, Ansible, Terraform, cloud-init, LVM, RAID, NVMe, NFS, Lustre, GPFS, SSH, TCP/IP, NVIDIA drivers, CUDA, container runtimes, DCGM, NVLink, NVSwitch, RDMA, InfiniBand, RoCEv2, PXE, Kickstart, libvirt, Qemu/KVM, SOC 2, and ISO 27001.