Staff Hardware Systems Engineer
Job description
About the role
You will own the full hardware lifecycle for Crusoe's production infrastructure, driving feasibility, bring-up, validation, and sustained reliability of high-performance compute systems from prototype through large-scale deployment. You will take end-to-end ownership of critical interconnects including PCIe, InfiniBand, and NVMe/storage, solving the most complex debugging and optimization challenges in the field. You will develop and maintain deep automation frameworks for hardware testing, diagnostics, and continuous reliability improvements that scale across our global infrastructure. Your work will directly accelerate the deployment of sustainable, AI-first compute systems and ensure world-class performance for our customers and partners. You will collaborate daily with cross-functional teams across energy, manufacturing, data center construction, firmware, software, and cloud services to resolve system-level issues and enable stable production operation. You will identify opportunities for new hardware technologies and testing methods that align with Crusoe's long-term mission to make energy and intelligence abundant for the world. If you want to do the most meaningful work of your career and help advance AI strategies for our customers and partners, this is your path to grow alongside a high-performing, expert team.
Key facts
What you'll do
- Drive the full hardware development and sustaining lifecycle, including feasibility, bring-up, validation, deployment, and ongoing production support for AI compute infrastructure.
- Develop and maintain scripting and automation frameworks for hardware testing, diagnostics, and continuous reliability improvements across large-scale deployments.
- Lead deep troubleshooting and debugging for PCIe (link training, topology, performance issues), InfiniBand (fabric debugging, throughput, connectivity issues), and NVMe/storage (performance bottlenecks, firmware interactions, failure analysis).
- Conduct rigorous system validation and characterization for GPU, CPU, and high-performance compute platforms under real-world and edge-case conditions.
- Support end-to-end integration and solution testing to ensure Crusoe Cloud products meet performance, reliability, and scalability expectations for demanding AI workloads.
- Collaborate with mechanical, thermal, firmware, software, and manufacturing teams to resolve system-level issues and enable stable, repeatable production operation at scale.
- Drive prototyping, qualification, and readiness for high-volume manufacturing with both internal engineering teams and external vendors to accelerate timelines.
- Identify opportunities for new hardware technologies, testing methods, and sustainability improvements that support Crusoe's long-term energy and intelligence abundance mission.
- Provide data-driven insights and evidence-based recommendations to influence Crusoe's hardware roadmap and reliability strategy across product and infrastructure teams.
- Partner with cloud and platform teams to validate hardware features, enable firmware and driver optimizations, and support new product introductions.
- Monitor field failure trends, drive root cause analysis, and implement corrective actions to improve yield, availability, and customer satisfaction.
- Maintain and evolve test plans, test procedures, and validation documentation to ensure compliance, traceability, and continuous improvement.
- Mentor and collaborate with engineers across disciplines to elevate hardware practices, share knowledge, and reduce time-to-resolution for complex issues.
- Track and report on key hardware metrics, including performance, reliability, power, and thermal characteristics, to guide investment and optimization decisions.
- Contribute to the design for testability and debuggability of future hardware platforms, ensuring that production systems are observable and maintainable.
Requirements
- 8-10+ years of experience in hardware development, validation, sustaining engineering, or production engineering within compute, cloud, or high-performance environments.
- Strong hands-on expertise in PCIe (link training, error detection, performance tuning), InfiniBand (fabric configuration, routing, performance analysis), and NVMe/storage (controller firmware, performance profiling, failure modes).
- Deep proficiency in hardware bring-up, board-level debugging, and system-level validation for complex server architectures and GPU-centric workloads.
- Ability to design and implement automation frameworks for hardware testing, diagnostics, and continuous reliability improvements using scripting languages such as Python and Shell.
- Technical background in digital and analog design, server architecture, and high-performance compute hardware, including power, thermal, and signal integrity considerations.
- Experience working across thermal, mechanical, firmware, software, and manufacturing functions in multidisciplinary, fast-paced environments.
- Strong analytical and problem-solving skills with a data-driven approach to identifying root causes and quantifying impact in production systems.
- Excellent communication and collaboration skills for working effectively with internal teams and external partners across geographies and time zones.
- Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, or equivalent hands-on experience that demonstrates the required technical depth.
- Proven track record of owning complex hardware issues from initial symptom analysis through resolution in production-scale deployments.
- Demonstrated ability to read and interpret hardware schematics, signal integrity analyses, and specification documents to drive debug and validation activities.
- Experience with version control, continuous integration concepts, and infrastructure-as-code practices relevant to hardware validation and test automation.
- Willingness to travel as needed to support manufacturing validation, customer deployments, and field investigations, and to work in US time zones to support global operations.
Nice to have
- Experience designing or optimizing GPU-to-GPU communication architectures for AI/ML workloads to improve scalability and reduce latency.
- Direct experience integrating NVLink or other next-generation GPU interconnect technologies in production systems.
- Familiarity with cutting-edge GPU architectures and how to leverage them in AI/HPC environments for performance and efficiency optimization.
- Expertise supporting or designing systems across both ARM and x86 server architectures to enable flexible deployment options.
- Background in sustainable or energy-efficient hardware design practices that align with Crusoe's mission to maximize the abundance of energy and intelligence.
- Advanced certifications or coursework in AI/HPC hardware systems, server architecture, or related engineering disciplines that deepen technical impact.
Practical notes
This role may require travel to support manufacturing validation, customer deployments, and field investigations as needed. You must be able to work in US time zones to coordinate with global teams and partners effectively.