Principal Infrastructure Engineer
Job description
About the role
Onebrief is building a new class of collaboration and AI-powered workflow software for military planning and operational coordination. The role is focused on owning the infrastructure platform that delivers our software to military staffs across a diverse and demanding operational landscape. You will be responsible for designing, building, and running the platform end to end, including Kubernetes clusters in commercial and government cloud, application streaming for thin-client users on classified networks, a unified artifact pipeline, and air-gapped appliances at forward sites. This position requires a bias towards action, where you solve complex infrastructure problems by weighing tradeoffs and moving forward with resilience. You will partner with Governance, Risk, and Compliance teams on STIGs, CVE remediation, audit prep, and the controls necessary to operate in environments such as IL5, IL6, and JWICS. Success is defined by a platform that remains stable under real operational load, ships trusted artifacts to every profile we run, and passes audits without fire drills.
Key facts
What you'll do
- Design, develop, test, iterate, and deploy secure production systems across cloud-native and edge appliance deployments.
- Embed with cross-functional teams and advise on infrastructure, security, and deployment best practices that hold up in production environments.
- Own end-to-end infrastructure outcomes for one or more critical programs or priority infrastructure initiatives.
- Harden the artifact pipeline so the same signed builds run in cloud, single-tenant high-side, and air-gapped appliance deployments.
- Build and operate the Kubernetes clusters that deliver our services, focusing on operators, networking, storage, and multi-cluster operations.
- Develop application streaming capabilities that enable thin-client users on classified networks to access our software securely and reliably.
- Create unified packaging and signing processes that ensure artifacts are consistent, verifiable, and deployable across every environment.
- Collaborate with mission owners and GRC teams to ensure deployments meet strict compliance requirements and are audit-ready at all times.
- Implement identity, network segmentation, and secrets management controls that align with security best practices and threat models.
- Troubleshoot and resolve complex infrastructure failures, drawing on deep production experience to maintain system resilience.
- Mentor engineers on infrastructure patterns, security fundamentals, and operating in regulated and high-stakes environments.
- Contribute to RFCs, PRDs, and production code, ensuring that infrastructure decisions are well-documented and technically sound.
- Partner with product and engineering teams to define and deliver infrastructure capabilities that unblock critical workflows.
- Continuously evaluate and adopt new technologies that improve reliability, security, and operational efficiency across the platform.
Requirements
- 8+ years operating production infrastructure in DevOps, DevSecOps, Platform, SRE, Cloud, or Infrastructure roles.
- Full-stack engineering experience when necessary, including designing, building, and operating services in a modern backend language such as Go, Python, Rust, or TypeScript.
- Deep production Kubernetes experience, including operators, networking, storage, multi-cluster operations, and an understanding of what breaks at scale.
- Strong expertise with major cloud providers like Amazon AWS and Microsoft Azure, with the ability to design across providers and operational profiles.
- Solid application and infrastructure security fundamentals, including identity, network segmentation, secrets management, common vulnerability classes, and sound security judgment under ambiguity.
- Demonstrated ability to build working solutions from scratch, connect disparate applications, and add value by diving into existing codebases.
- Experience packaging and shipping workloads into disconnected or air-gapped environments, ensuring integrity and consistency without internet access.
- Demonstrated leadership or mentorship experience, guiding engineers through complex technical challenges and operational issues.
- Ability to obtain and maintain a SECRET clearance, ensuring you can work on sensitive programs and data.
- Comfort operating in fast-paced, mission-critical environments where decisions have real-world consequences and resilience is essential.
Nice to have qualifications
- Active US SECRET or TOP SECRET security clearance.
- Government cloud expertise, including AWS GovCloud or Azure Government.
- Exposure to FedRAMP, IL4, IL5, IL6, or JWICS environments and the associated control frameworks.
- Audit experience with FedRAMP, SOC 2, or equivalent compliance programs.
- Big data and distributed data experience in high-throughput or regulated settings.
- Prior experience supporting defense, intelligence, or other regulated industries with strict operational requirements.
Practical notes
This is a full-time position based in the United States with remote work options. The role may require travel to customer sites or operational environments where decisions carry significant consequences. Visa sponsorship may be considered for exceptional candidates where legally permitted. Applicants must be able to obtain a SECRET clearance to proceed in the hiring process.