Platform Engineer (Forward Deployment)
Job description
About the role
Lightning AI is seeking experienced engineers to act as the primary technical partner for our customers. You will architect, build, and deploy production-grade AI systems while translating complex requirements into scalable platform solutions. In this capacity, you will serve as the definitive technical authority for client engagements, ensuring that every solution aligns with both business goals and architectural best practices. You will own the design decisions that bridge the gap between ambitious AI concepts and reliable, production-ready implementations. The role demands a unique blend of deep engineering skill and consultative presence to guide customers through their most critical deployment challenges. You will be responsible for converting abstract ambitions into concrete, maintainable software that delivers measurable value. Ultimately, your work will define the reliability and performance profile of the Lightning platform in the hands of our most demanding users.
Key facts
What you'll do
- Partner with customers to design and implement end-to-end AI workflows on the Lightning platform, ensuring alignment with long-term operational goals.
- Convert high-level customer objectives into technical specifications and production deployments that are robust, scalable, and secure.
- Manage the full lifecycle of technical engagements, from initial discovery and scoping through to monitoring, optimization, and continuous scaling.
- Develop production-ready software and services, prioritizing Python to deliver efficient and maintainable codebases.
- Build observable systems with a focus on latency, throughput, and cost efficiency, enabling data-driven decisions on performance improvements.
- Debug and optimize AI systems, including model behavior, APIs, and distributed workloads, to resolve complex issues in real time.
- Collaborate with internal product and engineering teams to synthesize customer feedback and influence the platform roadmap.
- Make technical trade-offs to solve ambiguous problems without adding unnecessary complexity, preserving simplicity and maintainability.
- Conduct hands-on prototyping to validate technical hypotheses and de-risk challenging integration scenarios for clients.
- Serve as a mentor and thought partner, elevating the technical capabilities of client teams through guidance and knowledge transfer.
- Implement infrastructure patterns that enhance reliability, security, and compliance for deployments in regulated environments.
- Drive the adoption of best practices for version control, testing, and deployment automation across customer projects.
- Evaluate and integrate emerging AI tools and frameworks to ensure the platform remains at the forefront of technical innovation.
- Document all architectural decisions and implementation details to ensure continuity and clarity for ongoing support.
Requirements
- Professional experience building full-stack applications using modern frontend frameworks like React and TypeScript, demonstrating a strong grasp of component design and state management.
- Backend development experience using Python, Go, or similar languages, with a proven track record of writing clean, performant, and testable code.
- Direct experience working with customers in technical roles such as Solutions Engineering, Forward Deployed Engineering, or Applied AI, showcasing comfort in client-facing environments.
- Understanding of AI/ML pipelines, including model development, evaluation, and deployment, along with the associated operational considerations.
- Experience operating production AI/ML systems in cloud or distributed environments, with a deep appreciation for issues like resilience, monitoring, and incident response.
- Proficiency with AI infrastructure tools including Docker, Kubernetes, APIs, and model serving, enabling effective management of containerized workloads.
- Bachelor degree in Computer Science, Engineering, Mathematics, or a related field, providing a foundational understanding of computational principles.
- Demonstrated ability to work independently and collaboratively in fast-paced, dynamic settings where priorities evolve rapidly.
- Strong written and verbal communication skills, essential for translating technical concepts to diverse stakeholders.
- Willingness to travel occasionally to office hubs as required to maintain close collaboration with internal and external teams.
Nice to have
- Experience optimizing large-scale AI/ML systems in production, with measurable improvements in efficiency or cost reduction.
- Familiarity with AI stacks such as vLLM, TensorRT, Ray, LangGraph, or vector databases, providing insight into advanced deployment patterns.
- Knowledge of inference optimization, distributed systems, or GPU-accelerated workloads, allowing for deeper engagement with high-performance computing challenges.
- Previous experience in startup environments, where agility, ownership, and rapid iteration are highly valued.
- Advanced degree (Master or PhD) in a technical field, indicating a strong capacity for complex problem-solving and research-oriented thinking.
Practical notes
- This role requires a minimum of 2 days per week in one of our office hubs.
- Total rewards include base salary, discretionary bonus, and equity (RSUs).
- Benefits include health coverage, 401(k) matching (US) or pension contributions (UK), unlimited PTO, winter break, parental leave, professional development allowance, wellness stipends, and a 4-week sabbatical after four years.
- Complimentary meals are provided at office locations.