Member of Technical Staff
Job description
Member of Technical Staff at Emerald Ai.
About the role
Emerald AI is building the world's first power flexible managed cloud infrastructure and this role owns the end to end architecture and implementation of that managed cloud. The hire will architect platform services from zero to production while defining isolation boundaries, tenant models, and provisioning flows that scale across diverse infrastructure providers. You will engineer the control plane, self service interfaces, and automated lifecycle systems including usage metering tightly integrated with billing. The role requires deep collaboration with infrastructure partners to evaluate fabric quality, network isolation, and economics for bare metal GPU environments. You will design and implement rigorous multi tenancy across compute, storage, and networking ensuring security, QoS, and encryption even for customers with root level access. The position drives workload orchestration across Kubernetes and Slurm for massive training and inference workloads while overseeing node health and driver management. You will lead the high performance storage strategy integrating parallel filesystems such as Lustre and ensuring seamless provisioning and operational behavior.
Key facts
What you'll do
Architect managed services from zero to production defining productization of GPU capacity, isolation boundaries, tenant models, provisioning flows, and service catalogs.
Engineer the platform core building robust control plane services, self service customer experiences, and automated lifecycle systems with integrated metering and billing.
Conduct deep technical assessments of bare metal GPU vendors evaluating fabric quality, network isolation, and economics to automate handoff to active tenants.
Implement end to end multi tenancy with strict isolation across compute, storage, and networking including InfiniBand and VLANs while enforcing security and encryption.
Drive workload orchestration across Kubernetes and Slurm environments managing node health, driver fleets, and kernel configurations for heterogeneous cloud deployments.
Deploy and integrate high performance parallel storage solutions such as Lustre, VAST, or Weka ensuring clean integration into the provisioning model.
Define and enforce SLOs, observability standards, and incident response protocols aligning internal practices with provider SLAs to deliver a reliable sellable product.
Evaluate, onboard, and vet infrastructure partners through hands on technical reviews covering performance, isolation, and cost effective scaling.
Design and implement systems for usage metering, quota management, and integration with billing infrastructure to support monetization of the platform.
Champion operational excellence by establishing monitoring, alerting, and debugging tooling that spans both our software stack and underlying provider environments.
Lead automation of deployment, scaling, and recovery workflows for GPU clusters ensuring high availability and performance at scale.
Collaborate closely with product and engineering teams to translate customer requirements into robust infrastructure primitives and APIs.
Own the design of networking, security, and storage layers to support diverse AI workloads with varying performance and isolation demands.
Guide best practices for GPU software stacks, including driver compatibility, firmware management, and health diagnostics across distributed nodes.
Requirements
Bring at least 7 years of infrastructure or platform engineering experience with a track record of launching managed cloud or AI platforms that served production users.
Demonstrate strong hands on experience with Kubernetes and Slurm and proven ability to offer them as managed services with defined SLAs.
Show production experience deploying or operating parallel filesystems such as Lustre, GPFS, Weka, VAST, or BeeGFS including deep knowledge of architecture, tuning, and failure modes.
Possess a comprehensive understanding of cloud service fundamentals including control planes, tenancy and isolation models, APIs, quota and metering systems, and the operational discipline required for revenue facing services.
Exhibit mature infrastructure as code practices using tools such as Terraform and Ansible and have solid programming ability in Python or Go.
Show familiarity with GPU infrastructure including high performance networking with InfiniBand, RoCE, and RDMA as well as the GPU software stack and driver lifecycle.
Have a strong grasp of Linux systems administration, performance troubleshooting, and security hardening across distributed compute and storage environments.
Demonstrate the ability to design systems for multi tenant isolation, secure data boundaries, and compliance requirements in shared infrastructure.
Display excellent communication skills to collaborate with cross functional teams, partners, and customers while documenting designs and operational procedures.
Nice to have
Prior experience at a GPU cloud, hyperscaler AI service, or HPC center delivering compute and storage as a service, particularly one built on rented or colocated capacity.
Familiarity with NVIDIA reference architectures such as SuperPOD, GPUDirect Storage, NCCL debugging, and DCGM metrics and diagnostics.
Hands on experience with Lustre multitenancy features including nodemap, fileset mounts, and Kerberos, or with service provider deployments of VAST or Weka.
Experience negotiating and integrating multiple infrastructure vendors and designing systems for portability between them.
Background running object storage at scale with systems such as S3, Ceph, or MinIO including thoughtful data tiering designs.
Experience building billing, metering, or FinOps pipelines for services that charge based on utilization and performance tiers.
Practical notes
The role is full time based in the Bay Area with no travel requirements and no visa sponsorship at this time.