Senior Software Engineer, Developer Experience
Job description
About the role
You will build tools and workflows that improve how engineers interact with our AI-native cloud infrastructure. This role focuses on creating internal platforms that increase productivity and simplify the deployment of complex workloads across our GPU-accelerated environment. You will own the design of internal developer platforms that abstract the complexity of our infrastructure for end users. The work you ship will directly influence the velocity and satisfaction of engineering teams building AI workloads. You will partner closely with infrastructure and platform teams to ensure the tools you create are robust and scalable. This position requires a mindset focused on usability, reliability, and the end-to-end developer journey. You will be responsible for turning operational complexity into simple, self-service experiences.
Key facts
What you'll do
- Design and implement internal developer tools to streamline infrastructure management and reduce manual overhead.
- Improve the interface between our hardware resources and the software teams using them by creating intuitive control planes.
- Build automation for fleet and node lifecycle management, ensuring consistency and reliability at scale.
- Create documentation and self-service portals to reduce operational friction and empower developers to solve their own problems.
- Collaborate with infrastructure teams to optimize the developer loop for AI training and inference workflows.
- Instrument and measure the developer experience to identify bottlenecks and drive data-informed improvements.
- Develop robust APIs and CLIs that enable engineers to interact with our platforms programmatically and efficiently.
- Implement observability into developer-facing services to ensure reliability and fast troubleshooting of issues.
- Partner with product teams to translate business requirements into scalable infrastructure solutions.
- Refactor legacy tooling to improve performance, maintainability, and alignment with current best practices.
- Act as a technical advocate for developers, ensuring their feedback shapes the direction of the platform.
- Contribute to open source projects and internal standards that enhance the overall ecosystem of tools.
Requirements
- Professional experience in software engineering with a focus on platform or developer tooling accumulated over your career.
- Proficiency in building scalable systems for cloud environments where performance and uptime are critical.
- Ability to write clean, maintainable code in languages commonly used for infrastructure automation and scripting.
- Experience working with Kubernetes and container orchestration at scale to manage complex distributed workloads.
- Understanding of distributed systems and cloud-native architecture principles that guide resilient system design.
- Strong problem-solving skills with the ability to debug complex issues in distributed and multi-component systems.
- Excellent written and verbal communication skills to collaborate effectively with cross-functional teams.
- A proactive approach to learning new technologies and translating them into solutions for internal users.
- Commitment to writing tests and ensuring the reliability of the tools you deliver to other engineers.
- Experience with infrastructure as code practices and version control workflows using modern tooling.
- Ability to work asynchronously and maintain context across distributed teams and time zones.
- Willingness to dive into unfamiliar codebases and environments to diagnose issues and implement fixes.
Nice to have
- Background in high-performance computing or GPU-accelerated environments to better understand the workloads being supported.
- Experience with observability tools and telemetry systems to build insights into platform usage and performance.
- Familiarity with the lifecycle management of large-scale hardware clusters and the challenges they introduce.
- Knowledge of networking and storage systems as they relate to AI workloads and data movement.
- Experience contributing to or maintaining internal platforms that serve a large and diverse set of users.
- Understanding of security and compliance considerations in shared multi-tenant cloud environments.
Practical notes
CoreWeave is an AI-native cloud provider focused on high-performance GPU compute. Applicants should be prepared to work from one of the listed office locations. Full-time engagement is expected, with standard working hours aligned with the role. No specific visa sponsorship details or deadlines are provided within the source material. Travel is not mentioned as a requirement for this position. This role is based in specific geographic locations and remote work arrangements are not detailed in the available source information.