Senior Principal Platform Engineer
Job description
About the role
National Security Platform Engineering Opportunity in Jessup, MD
Clarity Innovations operates as a national security partner focused on safeguarding United States interests. The organization delivers innovative solutions that empower the Intelligence Community and the Department of Defense to transform data into actionable intelligence. This mission ensures success within an evolving operational environment. The company specializes in modernizing data operations through advanced workflows, CI/CD, and secure DevSecOps practices. Current challenges in Information Warfare, Cyber Operations, Operational Security, and Data Structuring define the scope of their end-to-end solutions.
The company seeks a Platform Engineering Architect to operationalize, secure, and mature artificial intelligence capabilities. This role focuses on building and maintaining the AI "path-to-prod" and a scalable AI run stack, which functions as the AI plumbing. Success requires providing a unified interface for AI model access. The position integrates foundational and reasoning models, AI agents, and generative AI.
Role Responsibilities
A core duty involves contributing to the Kubernetes-based AI access platform. The architect will manage the deployment of core AI services. Integration of frontier models, such as Claude and GPT, along with local inference engines like vLLM, is essential. Designing MLOps pipelines connects vLLM and Ray with MLflow to enable reliable model lifecycle management. The role provides expert recommendations for enterprise AI adoption. Specific guidance covers agentic orchestration and spec-driven development.
Professional Requirements
The position mandates professional experience spanning 7+ years in combined DevSecOps, Platform Engineering, or SRE. Mastery of Kubernetes is required, including Administration and Development of clusters. Advanced knowledge of containerization, specifically Docker or equivalent build tools, is mandatory. Experience with Cloud and Infrastructure as Code, such as Azure or AWS architecture and Terraform/Crossplane, is necessary.
Proficiency in at least one backend or scripting language is required. Python, Go, or Bash drive systems automation. A solid understanding of OSI Layer 4-7 is essential. This includes VPC/VNET configuration, DNS, Load Balancing, and SSL/TLS management. Practical experience with AI/ML frameworks and tooling defines the technical baseline. Relevant tools include PyTorch, Hugging Face, LangChain, vLLM, Ray, and MLflow.
Hands-on experience deploying, scaling, and securing AI/ML workloads on Kubernetes is critical. This includes GPU-enabled clusters, model-serving platforms, and distributed inference/training systems. Experience building internal AI platforms or developer enablement tooling supports model lifecycle management. Familiarity with MLOps concepts covers automated model deployment, versioning, evaluation, observability, rollback strategies, and CI/CD integration for AI systems. Working knowledge of vector databases, embedding pipelines, retrieval-augmented generation (RAG), and semantic search architectures completes the requirements.
The role establishes secure DevSecOps workflows. These workflows ensure resilient data operations for cyber and information warfare missions. Implementing vector and retrieval systems strengthens RAG and semantic search for structured data workflows.
What you'll do
- Architect and own the design of the national security platform engineering solutions, ensuring alignment with Defense and Intelligence community standards.
- Drive the development and lifecycle management of the AI platform, focusing on production-grade reliability, scalability, and security for classified workloads.
- Lead the integration of large language models and generative AI tools, including Claude, GPT, and open-source frontier models, into a cohesive runtime environment.
- Engineer and maintain the AI run stack, encompassing vLLM, Ray, and MLflow, to support high-throughput inference and distributed training operations.
- Define and implement MLOps pipelines that automate model deployment, version control, evaluation, and rollback capabilities for AI systems in regulated environments.
- Provide technical leadership and expert consultation on agentic orchestration frameworks and spec-driven development practices for AI-powered applications.
- Configure and optimize VPC/VNET, load balancing, DNS, and SSL/TLS infrastructure to meet the stringent performance and security demands of national security missions.
- Collaborate closely with cyber and information warfare teams to establish secure DevSecOps workflows that protect data integrity and operational resilience.
- Develop and govern vector databases, embedding pipelines, and retrieval-augmented generation architectures to enable semantic search and structured data workflows.
- Establish platform observability, monitoring, and logging standards to ensure end-to-end visibility into AI model performance and infrastructure health.
- Mentor platform engineers and developers, elevating the team's capability in Kubernetes administration, container security, and AI workload optimization.
- Partner with mission owners to translate operational requirements into scalable platform capabilities that support rapid prototyping and hardened deployment.
- Champion secure coding practices and infrastructure-as-code methodologies to accelerate delivery while maintaining compliance with defense controls.
- Evaluate emerging AI infrastructure tools and frameworks, conducting proof-of-concept assessments to identify capabilities that enhance mission outcomes.
- Ensure all platform activities adhere to strict operational security protocols and data handling regulations governing national security systems.
Requirements
- Possess a minimum of 7 years of cumulative experience in DevSecOps, Platform Engineering, or Site Reliability Engineering roles within high-assurance environments.
- Demonstrate expert-level proficiency in Kubernetes, including cluster installation, configuration, networking policies, and daily administration.
- Exhibit advanced competency in containerization technologies, with mastery of Docker, containerd, and associated build and runtime tools.
- Hold experience designing and managing infrastructure through code using Azure, AWS, and configuration management tools like Terraform and Crossplane.
- Show fluency in at least one systems or scripting language, such as Python, Go, or Bash, to develop automation and tooling.
- Display a thorough comprehension of networking models defined by the OSI stack, specifically layers 4 through 7, including load balancers, firewalls, and encryption protocols.
- Apply deep knowledge of AI and machine learning frameworks, including PyTorch, Hugging Face Transformers, LangChain, and model serving ecosystems.
- Have a record of deploying and operating AI and ML workloads at scale on Kubernetes, including GPU cluster management and inference services.
- Understand MLOps end-to-end, covering data versioning, model evaluation, monitoring, rollback procedures, and CI/CD integration tailored for regulated AI systems.
- Hold experience with vector storage, embedding techniques, RAG implementation, and semantic search architectures used in secure data platforms.
Nice to have
The role prefers candidates who have built and operated AI platforms for enterprise or defense applications. Experience with classified or high-security environments is valued. Backgrounds in information warfare, cyber defense operations, or data structuring for national security missions are advantageous. Familiarity with defense compliance frameworks, such as NIST or RMF, supports secure delivery.
About the company
Clarity Innovations delivers software, data, and cyber solutions for Department of Defense and federal missions. Teams design, build, and deploy systems that operate in secure, scalable environments. Work focuses on clear requirements, modular architectures, and measurable outcomes. The culture values disciplined engineering, candid communication, and steady delivery that keeps projects moving at operational speed.