Systems Engineer, VDI Platform
Job description
About the role
CoreWeave is building the foundational cloud infrastructure for AI and is seeking a Systems Engineer to own the end-to-end delivery of a next-generation VDI platform. This role is the primary technical owner for the remote-compute infrastructure that will power hundreds of engineers and support teams across the company. You will design, build, and operate the full stack from cloud infrastructure to OS image lifecycle for both Ubuntu and Windows 11 virtual desktops. The position requires deep collaboration with security, identity, and operations teams to deliver a platform that is secure, self-service, and operationally excellent. You will be responsible for every phase of the platform lifecycle, from initial architecture and design through steady-state operations and continuous improvement. This is a hands-on role that demands ownership, rigor, and a commitment to reliability as CoreWeave scales globally.
Key facts
What you'll do
Design and deliver the end-to-end VDI platform architecture, covering cloud infrastructure, access controls, OS lifecycle, security posture, and network topology, and drive peer review and sign-off before build commences.
Integrate cloud infrastructure into CoreWeave's Teleport cluster, including IAM node joining, RBAC role definitions, access policies per user persona, and break-glass/OOB node setup and validation.
Configure Okta SAML/OIDC SSO with MFA enforcement, session policy, and access review workflows for the VDI fleet.
Build and maintain the audit logging pipeline, directing cloud logging into CoreWeave SIEM, and manage legal hold and forensic snapshot-on-demand workflows alongside compliance-aligned session activity retention policies.
Own cost governance by implementing instance tagging and cost center chargeback attribution, building idle detection and cleanup automation, and delivering self-service and admin portals for instance lifecycle management.
Own the Ubuntu LTS and Windows 11 base image pipelines using code-defined builds, automated regression test suites covering agent health, network reachability, and security posture, versioned release promotion, and rollback procedures.
Write and maintain Chef cookbooks or equivalent configuration management for post-provision setup, including user setup, mounts, and toolchain configuration.
Bake security and network agents such as CrowdStrike, Netskope, and BlastShield, along with the PCoIP client and agent, into both Ubuntu and Windows 11 base images, and own fleet enrollment and alert routing.
Lead end-to-end integration testing from provision through authentication to tool access and teardown, and conduct platform security reviews prior to launch, driving remediation and sign-off.
Sustain steady-state operations through monthly patch releases, configuration management updates, infrastructure health checks, incident response, and user ticket triage.
Produce and maintain comprehensive runbooks covering image release procedures, Teleport onboarding, break-glass recovery processes, and incident response playbooks.
Partner closely with security and identity teams to ensure access policies, session governance, and compliance requirements are met and continuously validated.
Continuously monitor platform performance and user experience, driving optimizations for reliability, scalability, and operational efficiency.
Act as the day-one expert on the VDI platform, serving as the central point of ownership for roadmap execution, technical decisions, and cross-functional coordination.
Requirements
Demonstrated experience operating and automating cloud infrastructure on AWS, Azure, or GCP, with a strong grasp of compute, storage, and networking primitives.
Proficiency with infrastructure-as-code tools such as Terraform or Pulumi, and configuration management using Chef or similar declarative frameworks.
Solid understanding of identity and access management concepts, including SAML and OIDC, with hands-on experience implementing SSO in enterprise environments.
Experience designing and operating audit logging pipelines, log analytics, and compliance-driven retention policies in complex cloud environments.
Familiarity with VM lifecycle management, image building methodologies, and patching strategies for both Linux and Windows workloads.
Working knowledge of networking and security tooling common in enterprise environments, including firewalls, load balancers, and endpoint protection platforms.
Experience with monitoring, observability, and incident response in distributed, cloud-native systems.
Strong written and verbal communication skills, with the ability to document designs, runbooks, and operational procedures for technical and non-technical audiences.
Practical notes
This role is based in Livingston, New Jersey, with the flexibility to work from New York, Sunnyvale, California, San Francisco, California, Bellevue, Washington, or remote locations as determined by CoreWeave. The position requires collaboration across multiple time zones and on-call responsibilities as part of the operational model. CoreWeave is an equal opportunity employer and is committed to building a diverse and inclusive workforce. All employment decisions are made on the basis of required skills and qualifications.