Staff Software Engineer, Network Development
Job description
About the role
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. You will own the design, development, and long-term health of the platforms and services that automate the network, including config generation, device provisioning, workflow automation, observability tooling, and internal CLIs and APIs. You will plan and map out complex projects months in advance, sequencing the work, anticipating dependencies, and keeping several efforts moving in parallel. You will lead technical direction through RFCs, design reviews, and architecture decisions, and break ambiguous problems down into shippable work. You will write well-tested, well-documented software in Python and Go, and set the standard for quality across the team through strong test cases, clear error handling, and solid CI/CD practices. You will improve the reliability and operability of what we run by defining service-level indicators and objectives and relentlessly reducing manual toil through automation. You will partner closely with Network Engineering, Observability, Fleet Engineering, and HPC teams so our automation fits cleanly into the wider system. You will join the on-call rotation once you are hired.
Key facts
What you'll do
Define the architecture for network automation platforms and services that scale with the world's largest GPU deployments.
Own the full lifecycle of production network software, from initial design and implementation through observability, maintenance, and iterative improvement.
Translate ambiguous operational requirements into concrete technical specifications and shippable deliverables.
Design and implement configuration generation and device provisioning systems that are robust, repeatable, and self-healing.
Build workflow automation and orchestration tools that reduce manual steps and increase network reliability.
Instrument network services with rich observability, defining service-level indicators and objectives to guide reliability efforts.
Collaborate with Network Engineering to translate low-level device behaviors into stable, automatable software abstractions.
Partner with Observability and SRE teams to ensure telemetry, alerting, and dashboards support fast, data-driven decisions.
Work alongside Fleet Engineering and HPC teams to ensure network automation integrates seamlessly with compute and storage systems.
Contribute to internal CLIs and APIs that enable other teams to interact with network services safely and efficiently.
Write high-quality Python and Go code, with comprehensive tests, clear documentation, and reliable CI/CD pipelines.
Lead technical discussions through RFCs and design reviews, breaking complex problems into manageable, shippable work.
Drive standards and best practices for software quality, error handling, and operational excellence across the network team.
Participate in on-call rotation to respond to incidents and support continuous service reliability.
Continuously explore new tools and techniques to reduce manual toil and improve the operability of network infrastructure.
Requirements
Demonstrate strong software engineering fundamentals and the ability to design and implement systems that are maintainable and scalable.
Bring a solid understanding of networking fundamentals, including routing, switching, and transport protocols relevant to data center networks.
Show experience writing production-grade software in Python and Go, with a proven track record of writing clean, testable, and well-documented code.
Exhibit comfort working with network devices through CLI, API, and NETCONF/YANG models in a data center environment.
Possess strong problem-solving skills and the ability to break down complex, ambiguous requirements into clear technical tasks and deliverables.
Have experience collaborating with cross-functional teams such as Network Engineering, Observability, SRE, and HPC in fast-paced, production environments.
Show commitment to reliability and operational excellence, including defining and using service-level indicators and objectives to guide automation and monitoring.
Be comfortable working in a fast-paced environment where priorities shift and ambiguity is common, and be proactive about driving clarity and execution.
Nice to have
Experience with large-scale network automation and observability platforms.
Familiarity with cloud provider network architectures and practices.
Contributions to open-source networking projects or relevant internal tools.
Experience with infrastructure-as-code and configuration management concepts relevant to network devices.
Practical notes
This role may require occasional travel to offices in Livingston, NJ; New York, NY; Sunnyvale, CA; and Bellevue, WA.
Employment eligibility to work in the United States is required for this position.
Please include your locations and eligibility to work in the United States in your application.
This role is eligible for standard paid time off and benefits as applicable to full-time employment at CoreWeave.