Systems Engineer, Kernel
Job description
About the role
CoreWeave is actively building the foundational layer for the next generation of artificial intelligence compute, and this role is central to that mission. The successful candidate will own the design, implementation, and optimization of the Linux kernel specifically for our hyperscale GPU cloud platforms. You will be responsible for ensuring the stability, performance, and efficiency of the core software stack that directly powers our customer workloads. This position requires a deep understanding of how hardware and software interact at the most fundamental levels to extract maximum performance. You will work in a fast-paced environment where your contributions will have a direct and immediate impact on the company's infrastructure. The role demands a high degree of ownership and the ability to solve complex, low-level problems that are critical to the business. You will be a key technical partner for our engineering teams, providing kernel-level expertise to guide architectural decisions.
Key facts
What you'll do
- Architect and implement kernel modules and subsystems to enable advanced GPU virtualization and direct hardware access for AI workloads.
- Diagnose and resolve intricate performance bottlenecks within the Linux kernel to maintain stringent service levels for compute-intensive applications.
- Conduct in-depth analysis of system traces and logs to identify root causes of instability and develop robust mitigation strategies.
- Partner with hardware engineers to develop and validate new device drivers that leverage the latest GPU and infrastructure silicon capabilities.
- Refactor existing kernel-level services to reduce latency and improve throughput in our distributed compute fabric.
- Evaluate upstream kernel community changes and backport or adapt them for our specific production environment and hardware stack.
- Implement security hardening measures at the kernel level to protect the integrity of our multi-tenant compute platforms.
- Automate the testing and validation of kernel configurations to ensure consistency and reliability across our global data center footprint.
- Collaborate with compiler and runtime teams to ensure optimal software toolchain integration with our kernel modifications.
- Document kernel interfaces and operational procedures to support the broader engineering organization and on-call responsibilities.
- Lead incident response efforts for critical kernel-related outages, guiding remediation and communication.
- Define and drive the execution of a long-term kernel roadmap that aligns with the company's aggressive product development goals.
- Mentor junior engineers on best practices for kernel development and debugging within a cloud-native context.
- Contribute to open source projects where relevant, balancing proprietary needs with community collaboration.
Requirements
- Demonstrate proficient expertise in C programming, with a strong track record of writing efficient and maintainable kernel code.
- Possess a deep understanding of Linux kernel internals, including process scheduling, memory management, and interrupt handling.
- Have substantial professional experience working with complex system architecture and low-level debugging in production environments.
- Show a proven ability to work on-site in one of our designated office locations at least four days per week.
- Exhibit mastery of device driver concepts and interaction between hardware components and the operating system.
- Bring experience with GPU compute architecture and the challenges of accelerating parallel workloads in a cloud context.
- Highlight a history of troubleshooting and resolving high-severity infrastructure issues under tight deadlines.
- Confirm eligibility to work in the United States without sponsorship for this role in the locations listed.
Nice to have
- Experience with high-performance networking stack optimizations relevant to distributed compute.
- Familiarity with virtualization technologies and hypervisor interactions with guest kernels.
- Contributions to mainline Linux kernel or other major open-source system projects.
Practical notes
Full-time
Must be able to work on-site in one of our designated office locations at least four days per week.
Candidates must be eligible to work in the United States without sponsorship for this role in the locations listed.
Applications will be reviewed on a rolling basis until the position is filled.