Not seeing the right opportunity right now?
Job description
Product Manager, Lightning AI
Lightning AI builds the platform that enables teams to develop, train, and deploy artificial intelligence systems. Our company emerged from the merger of Lightning AI and Voltage Park, uniting developer-first software with large-scale, cost-efficient compute. The result is an end-to-end environment where ideas move from research to production with reduced friction. We serve solo researchers, startups, and large enterprises, with offices in New York City, San Francisco, Seattle, and London. The company is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.
This page presents a specific role within that broader mission. The information that follows details a position focused on product delivery for the Lightning AI platform.
Who We Are
Lightning AI was founded in 2019 to accelerate the lifecycle of AI systems. The platform is built on PyTorch Lightning, an open-source project that simplifies machine learning engineering. Through the merger with Voltage Park, the company combines software tooling with dedicated compute resources. Security, observability, and control are integrated into every layer of the platform. The company operates globally, with teams contributing across time zones.
The Way We Work
People who thrive here are builders. They move quickly, communicate openly, take ownership, and continuously improve their work and their teams. Decisions are made with momentum, while thoughtful design ensures long-term scalability. The culture values action, honest conversation, and relentless improvement. Teams are empowered to own outcomes and build healthy, effective workflows. The focus is on creating systems that amplify human impact through AI and automation.
Not seeing the right opportunity right now?
Candidates who are excited by this description are encouraged to express interest even if a matching role is not currently published. This is an opportunity to be considered for future positions as they become available. Submitting information does not commit either party immediately. Staying connected through LinkedIn and monitoring the careers page is recommended for updates. Opportunities are available for US-based, UK-based, and remote candidates. Applications are managed through the official portal.
About the role
You own the full product journey from intake through build, review, and ship. Your work creates durable systems that let research ideas move into production with less friction. You communicate clearly, take ownership, and raise the bar for speed and quality. Collaboration across teams is central to the role, aligning engineering, research, and operations toward shared goals.
Key facts
Location options include London, England, United Kingdom; New York, New York, United States; Remote; and San Francisco, California, United States. Engagement is through the official application portal, with connections also encouraged via LinkedIn. The role is open to candidates based in the US and UK, as well as remote work arrangements. Compensation for this position ranges from 160,000 USD to 200,000 USD per year.
What you'll do
- Drive intake by turning ambiguous requests into clear product goals and user-focused requirements.
- Design and build core infrastructure components that support training and inference at scale.
- Run rigorous review sessions that validate performance, reliability, and security before release.
- Ship features using robust deployment patterns that keep systems observable and dependable.
- Forge partnerships by aligning with internal and external stakeholders on priorities and timelines.
- Instrument workflows with monitoring and debugging tools that clarify behavior and accelerate fixes.
- Challenge assumptions by questioning tradeoffs and proposing simpler, more maintainable solutions.
- Maintain documentation and runbooks that keep knowledge current and support consistent execution.
Requirements
You must have three to five years of experience building and deploying machine learning systems. A deep understanding of distributed training patterns and production inference constraints is essential. You communicate clearly and write precisely for both technical and nontechnical audiences. You are comfortable working in fast-moving settings where scope and priorities evolve. The ability to make decisions with incomplete information is a critical skill.
Nice to have
Experience contributing to open source projects like PyTorch Lightning is preferred. Familiarity with the project's architecture and contribution flow provides a strong foundation for the role.
Skills & tools
Proficiency with PyTorch Lightning, PyTorch, and Python is expected. Experience with Kubernetes, Docker, and major cloud platforms such as AWS, GCP, or Azure is part of the skill set for this position.
Practical notes
Information regarding location, compensation, and responsibilities is subject to change until formally posted. All decisions regarding hiring and compensation are made in accordance with company policies.