Staff Security Engineer, Infrastructure
Job description
Staff Security Engineer, Infrastructure at Fal Ai.
About the role
The role of Staff Security Engineer, Infrastructure at Fal Ai centers on protecting and advancing the security of the generative media ecosystem that powers AI products in production. You will own the design and execution of security controls across the full infrastructure stack that enables fal.ai's platform, including compute, networking, storage, and data pipelines. This position requires a systems-oriented mindset to secure GPU-heavy workloads, multi-cloud environments, and the identity and access mechanisms that connect them. You will operate at the intersection of infrastructure engineering and security, implementing practical controls that scale without degrading performance. The role is hands-on and collaborative, working directly with platform and product teams to embed security into the foundation of every service. You will contribute to the strategic direction of platform security while delivering tactical improvements that harden systems against evolving threats.
Key facts
What you'll do
Design and implement security controls across cloud infrastructure, ensuring that network segmentation, firewall policies, and encryption in transit and at rest are applied consistently.
Harden Kubernetes and containerized workloads by defining secure baselines, admission controls, and runtime protections for GPU and CPU-based services.
Manage machine identity and workload authentication, leveraging protocols and standards that enable secure communication between services without excessive friction.
Operate and maintain secrets management and encryption systems, including integration with KMS and Vault to protect sensitive keys and credentials.
Implement least-privilege access and short-lived credentials across infrastructure and application layers, reducing the attack surface for both external and insider threats.
Architect Zero Trust security models that span networking, service meshes, and edge systems, ensuring verification and minimal trust across all connections.
Protect model weights, inference endpoints, and customer data through secure data access pathways, isolation mechanisms, and tenant-aware enforcement.
Design secure multi-tenant execution environments that prevent cross-tenant interference while preserving the performance required for AI workloads.
Build security guardrails directly into infrastructure and CI/CD pipelines using Infrastructure-as-Code patterns to enforce secure defaults automatically.
Use Terraform and similar tools to codify security policies, ensuring that new environments are provisioned with secure configurations from the outset.
Continuously identify and remediate security gaps through automated scanning, monitoring, and response workflows integrated into operational processes.
Collaborate closely with platform, infrastructure, and machine learning teams to embed security into architecture reviews and design decisions.
Enable engineering teams to move quickly by providing secure-by-default platforms and clear guidance rather than manual approvals for every deployment.
Contribute to threat modeling exercises, risk assessments, and security initiatives that reduce exposure across the stack and improve resilience.
Drive projects focused on network isolation, encryption, secure service communication, and observability to detect and respond to incidents effectively.
Requirements
Bring 8 or more years of experience in security engineering, infrastructure, or SRE roles that involve production systems at scale.
Demonstrate a strong understanding of cloud security principles and practices across major providers such as AWS, GCP, or Azure.
Show deep familiarity with networking fundamentals, including segmentation, firewall design, and Zero Trust concepts applied to distributed systems.
Possess hands-on experience with Linux systems and container security, including Docker and Kubernetes, and how to secure orchestrated workloads.
Have a record of building or securing production infrastructure at scale, with evidence of operating in high-availability and high-throughput environments.
Exhibit deep knowledge of authentication and authorization systems, including protocols, identity providers, and access patterns for APIs and services.
Understand secrets management and cryptography basics, including how to protect data and keys throughout their lifecycle in distributed systems.
Show competence in identifying common vulnerabilities and attack vectors, and applying security controls to mitigate risks across infrastructure layers.
Possess strong engineering skills, including proficiency in at least one general-purpose language such as Go or Python for automation and tooling.
Have substantial experience with Infrastructure-as-Code, with a preference for Terraform, to define and enforce secure configurations programmatically.
Demonstrate a strong automation mindset, ensuring that security processes scale alongside infrastructure growth and complexity.
Nice to have
Prior experience with GPU infrastructure or ML systems that require specialized security considerations for model execution and data protection.
Background in multi-tenant platform isolation, including mechanisms for workload separation, resource governance, and tenant-aware security policies.
Exposure to service mesh and zero-trust architectures that enforce strict communication controls and observability between services.
Experience working in high-growth startup environments where security must evolve quickly alongside product velocity and operational demands.
Practical notes
This is a full-time position based in San Francisco. The role requires availability during standard business hours and on-call rotations as needed for incident response and infrastructure changes. Candidates must be authorized to work in the United States without sponsorship for this position. Relocation support is not provided. The posting will remain open until filled, and applications will be reviewed on a rolling basis.