Sr. AI Engineer, Platform Infrastructure, Special Programs
Job description
About the role
This position involves creating reliable, deterministic tooling for engineers operating in classified environments where remote support is unavailable. You will develop the infrastructure that powers our special programs, ensuring that systems are fully tested and ready for deployment across various security levels. The role requires a disciplined approach to automation in air-gapped contexts where standard DevOps practices must be adapted for strict compliance. You will own the entire lifecycle of platform components from initial design through decommissioning. Collaboration with cross-functional teams will be essential to translate complex operational requirements into robust software solutions. Success in this role depends on your ability to build systems that are both secure and scalable under constrained conditions. You will be responsible for maintaining the integrity and performance of critical infrastructure that supports sensitive mission objectives.
Key facts
What you'll do
- Design and implement a deployment generator that standardizes application rollouts across heterogeneous environments.
- Develop and maintain comprehensive documentation for the platform to ensure knowledge transfer and operational clarity.
- Construct the bundle pipeline that packages applications and dependencies for secure distribution.
- Execute profile-driven rendering for enterprise on-prem, public cloud, and air-gapped classified targets to meet specific operational needs.
- Manage the cross-profile CI matrix to ensure consistent builds and tests across all supported configurations.
- Create a validation and testing framework that verifies system behavior before deployment to sensitive environments.
- Integrate inputs from the Supercompute team, including CRDs, Helm charts, and operators, without altering production systems or disrupting existing workflows.
- Lead the migration of the monitoring stack to improve observability and system health insights.
- Manage the pipeline for bare metal provisioning to support infrastructure that cannot rely on virtualized resources.
- Implement infrastructure as code patterns that enable repeatable and auditable deployment processes.
- Coordinate with security teams to ensure all tooling complies with regulatory and organizational standards.
- Optimize resource utilization across compute and storage systems to support high-performance computing requirements.
- Troubleshoot complex infrastructure issues using logs, metrics, and direct system inspection.
- Mentor junior engineers on best practices for building resilient and secure platform components.
Requirements
- Bachelor degree in physics, mathematics, computer science, computer engineering, data science, or an engineering field.
- 5+ years of experience with Python and/or Go Lang.
- 5+ years of experience with Kubernetes or similar tools for scaling, developing, and automating containerized applications.
- Must be a U.S. citizen, lawful permanent resident, refugee, or asylee to meet ITAR requirements.
- Ability to secure and maintain a Top Secret or Top Secret SCI security clearance.
- Demonstrated experience working in environments with strict compliance and audit requirements.
- Strong understanding of networking concepts and distributed system communication patterns.
- Proven ability to write clean, maintainable, and well-documented code.
Nice to have
- Experience with CI/CD pipelines and deployment automation tools like GitHub Actions or Buildkite.
- Knowledge of container image management and registry operations.
- Proficiency in Linux systems and shell scripting for automation tasks.
- Experience with Infrastructure-as-Code tools such as Terraform or Pulumi, including understanding state management.
- Experience building disconnected or air-gapped deployment tools where all dependencies are pre-staged.
- Familiarity with GPU infrastructure, including NCCL, NVIDIA drivers, CUDA, and InfiniBand/RoCE networking.
- Experience with bare metal provisioning technologies like cloud-init, PXE boot, or squashfs.
Skills & tools
- Python, Go Lang, Kubernetes, CI/CD, Linux, Infrastructure-as-Code (Terraform/Pulumi), GPU infrastructure (NVIDIA/CUDA/NCCL), Bare metal provisioning.
Practical notes
-
Salary: $220,000.00 - $350,000.00 per year.
- Travel: Up to 20% travel to government sites may be required.
- Schedule: Willingness to work weekends and extended hours is required.
- Benefits: Includes company stock, stock options, long-term cash awards, discretionary bonuses, Employee Stock Purchase Plan, 401(k), medical, dental, vision, life insurance, short/long-term disability, paid parental leave, 3 weeks vacation, and 10+ paid holidays.
- Clearance: Employment is contingent upon obtaining and maintaining a Top Secret security clearance. Failure to obtain this clearance may result in termination.