Sr. Principal Platform Engineer
Job description
About the role
Clarity Innovations is a trusted national security partner, dedicated to safeguarding our nation's interests and delivering innovative solutions that empower the Intelligence Community (IC) and Department of Defense (DoD) to transform data into actionable intelligence, ensuring mission success in an evolving world. The Sr. Principal Platform Engineer will own the design and execution of resilient, scalable platforms that underpin critical mission workflows. You will architect and operate Kubernetes and cloud-native infrastructure that enables secure and efficient delivery of data operations for national security customers. This role demands a "Whatever It Takes" mindset, driving pragmatic execution and rapid delivery without compromising engineering rigor or security standards. You will lead infrastructure decisions that directly impact operational outcomes and mission impact. The ideal candidate thrives at the intersection of infrastructure, automation, and observability in high-stakes environments. You will partner closely with engineers and operators to ensure systems are reliable, performant, and aligned with evolving mission needs.
Key facts
What you'll do
- Lead the design and implementation of Kubernetes and cloud-based systems on AWS and Azure, aligning architecture with mission-critical requirements.
- Provision and manage infrastructure as code using Terraform, Crossplane, and Ansible to ensure repeatable, auditable, and secure environments.
- Design and operate specialized environments using tools like KIND to bridge development and production workflows effectively.
- Deploy and maintain AI workloads where appropriate to solve complex customer challenges and enhance operational capabilities.
- Identify and resolve performance bottlenecks by optimizing service latency, resource utilization, and system throughput across the stack.
- Deploy, manage, and maintain Kubernetes clusters across development, staging, and production environments with a focus on reliability.
- Monitor and optimize workloads using Helm and Kustomize to deploy microservices at scale while maintaining operational efficiency.
- Implement network policies, load balancing, and service discovery using Istio Service Mesh to enhance connectivity and security.
- Design and operate CI/CD pipelines to streamline code deployments and enable rapid, reliable delivery of features and fixes.
- Ensure infrastructure security by implementing FIPS encryption, Zero Trust principles, and robust identity and access management practices.
- Advocate for relentless automation to eliminate manual processes, reduce risk, and improve system reliability and maintainability.
- Implement and monitor system availability and security using Prometheus and Grafana to provide actionable insights and maintain situational awareness.
- Troubleshoot and resolve production incidents with urgency, driving root cause analysis and implementing long-term automated fixes.
- Manage recovery processes and runbooks to ensure continuity of operations and rapid restoration of services during disruptions.
- Collaborate with cross-functional teams to understand high-level operational and business objectives for technology solutions.
- Contribute actively to sprint planning, daily stand-ups, and retrospectives to drive continuous improvement in engineering practices.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field or equivalent experience.
- 8+ years of professional experience in software engineering, platform engineering, or infrastructure roles.
- 8+ years of hands-on experience with Kubernetes administration, operations, and cluster lifecycle management.
- 8+ years of experience designing, building, and operating infrastructure using code, including proficiency in Terraform, Ansible, or similar tools.
- 6+ years of experience with AWS or Azure cloud platforms, including core services, networking, and security models.
- Demonstrated history of implementing and managing CI/CD pipelines and DevOps practices in production environments.
- Strong working knowledge of container orchestration, networking, and microservices architecture patterns.
- Ability to obtain and maintain a U.S. Government security clearance, if required for the position.
Nice to have
- Experience with service mesh technologies such as Istio in production environments.
- Familiarity with data processing pipelines and architectures for national security or defense applications.
- Knowledge of secure DevSecOps frameworks and compliance regimes relevant to government systems.
Practical notes
- Location requires hybrid work arrangement with a minimum of 3 days in the Herndon, VA office.
- This role involves U.S. government contract work which may require U.S. citizenship.
- Employment is contingent upon the successful completion of applicable background checks and clearance requirements.
- Travel requirements, if any, will be defined in the official offer terms and conditions.
- Candidates must meet all eligibility criteria listed in the requirements section to be considered.