Software Engineer, Site Reliability
Job description
Software Engineer, Site Reliability at Fal Ai.
About the role
You are a seasoned SRE who keeps production infrastructure running at scale. You own the reliability and availability of customer-facing systems - from Kubernetes clusters to deployment pipelines to the networking layer that connects it all. You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better. You will partner closely with platform and product teams to ensure that generative media workloads perform reliably under demanding conditions. Your work will directly impact how developers and enterprises build and deploy AI media pipelines on fal. You will play a key role in shaping the operational maturity of a rapidly growing infrastructure platform.
Key facts
What you'll do
- Architect and maintain the Kubernetes infrastructure that powers fal's generative media services, handling cluster lifecycle, upgrades, networking, and multi-tenant isolation for diverse customer workloads.
- Design, implement, and evolve our CI/CD pipelines and deployment infrastructure to enable safe, rapid, and reliable delivery of platform changes and customer features.
- Harness artificial intelligence to an extreme level to automate analysis and resolution of production issues, driving improvements in software development speed, reliability, and maintainability across the stack.
- Construct and refine observability platforms by building dashboards, alerting rules, and anomaly detection mechanisms that surface issues before they affect customers.
- Define, document, and enforce service-level objectives, and architect robust incident response processes that streamline communication and resolution during outages.
- Optimize and manage networking, load balancing, and service mesh configurations to ensure secure, high-throughput, and low-latency connectivity for all platform components.
- Champion reliability improvements by automating operational runbooks and introducing chaos engineering practices that validate system resilience proactively.
- Collaborate closely with engineering teams to translate reliability requirements into infrastructure solutions that support scalable and secure media processing.
- Investigate complex operational failures, conduct post-incident reviews, and translate findings into preventative measures that reduce future risk.
- Mentor other engineers on best practices for operating production-grade infrastructure in a fast-paced, AI-driven environment.
Requirements
- Bring 5+ years of experience managing critical production systems and software development workflows in fast-paced, high-visibility environments.
- Demonstrate strong production experience setting up and operating Kubernetes at scale, leveraging infrastructure-as-code tools such as Terraform and configuration management with Ansible.
- Show deep knowledge of Linux networking, container networking technologies including CNI plugins, VXLAN, and BGP, as well as enterprise-grade DNS implementations.
- Highlight experience building and operating CI/CD systems and GitOps workflows, with hands-on work using tools such as FluxCD and ArgoCD in production settings.
- Prove proficiency in Python and at least one of Go or Bash for developing automation scripts, tooling, and operational utilities.
- Illustrate strong experience with logging, monitoring, and alerting ecosystems, including Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, and Datadog in complex distributed environments.
- Communicate effectively with both technical and non-technical stakeholders, driving technical decisions and alignment across multiple teams and product areas.
- Operate as a self-starter who executes quickly, takes full ownership of assigned services, and constantly seeks opportunities for process improvement and innovation.
Nice to have
- Demonstrate experience with managing GPU and AI/ML workloads in production, including scheduling, resource isolation, and performance tuning for intensive media inference jobs.
- Apply kernel-based monitoring and routing techniques using eBPF and XDP to gain deep visibility into system and network behavior.
- Utilize security tooling such as Falco, Coroot, and SIEM platforms to detect threats, ensure compliance, and strengthen platform defenses.
- Implement and operate bare metal Kubernetes networking solutions, including Cilium, Calico, and MetalLB, to support high-performance on-premise deployments.
- Administer distributed storage systems like Ceph and Longhorn, ensuring that persistent storage for media assets meets stringent performance and durability requirements.
Practical notes
- This role is based in San Francisco, CA, and is offered as a full-time employment position.
- Candidates must be authorized to work in the United States without sponsorship for this position.
- The listed location and engagement details reflect the constraints and expectations defined in the original job description.
What we offer at fal
- Interesting and challenging work that pushes the boundaries of generative media infrastructure.
- A lot of learning and growth opportunities as you operate at the intersection of AI and platform engineering.
- We are currently hiring in downtown San Francisco within dynamic and collaborative workspaces.
- Comprehensive health, dental, and vision insurance coverage for you and your dependents across the United States.
- Regular team events and offsites designed to strengthen collaboration and celebrate shared achievements.