Staff Infrastructure Software Engineer
Job description
About the role
You own the platform that enables rapid engineering delivery while maintaining robust security and observability. You will define the evolution of infrastructure for a global AI-powered customer experience platform. The role operates within a small, action-oriented team that values autonomy and clear accountability. You are responsible for architecting and maintaining the foundational systems that empower engineering teams to deliver features safely and efficiently. This position requires a deep commitment to infrastructure reliability, security, and performance in a demanding, AI-centric environment. You will act as a key technical leader, guiding standards and best practices for the entire engineering organization. Your work will directly impact the scalability and resilience of the products used by customers worldwide.
Key facts
What you'll do
- Architect and maintain the core infrastructure platforms that enable rapid and secure engineering delivery across multi-cloud environments.
- Partner with cross-functional engineering teams to design and build developer tools that enhance workflow efficiency and deployment safety.
- Ensure the unwavering reliability and performance of Kubernetes clusters and associated CI/CD pipelines under demanding production conditions.
- Establish and enforce comprehensive metrics, logging, analytics, and alerting frameworks to monitor system performance and security across all endpoints and applications.
- Create and evolve infrastructure-as-code blueprints and deployment tooling that promotes reuse, standardization, and consistency across diverse engineering teams.
- Automate complex operational and engineering tasks to redirect effort toward high-value strategic initiatives and innovation.
- Build and maintain robust machine learning infrastructure that empowers AI teams to train, test, and deploy models efficiently on large-scale datasets.
- Manage critical partner integrations that connect Cresta with external systems to streamline processes and ensure smooth transitions from conception to production deployment.
- Implement and govern security and compliance controls within the infrastructure stack to meet organizational and regulatory requirements.
- Optimize cloud resource utilization and costs while maintaining high standards of service availability and operational excellence.
- Lead incident response efforts and contribute to the development of post-incident reviews to drive continuous improvement.
- Mentor junior engineers on infrastructure principles and tooling, fostering a culture of knowledge sharing and technical growth.
- Collaborate with SRE and platform teams to define service level objectives and ensure alignment with business goals.
- Evaluate emerging technologies and integrate proven innovations to keep the infrastructure platform at the forefront of the industry.
Requirements
- Must possess a minimum of five years of hands-on experience operating infrastructure in production settings.
- Must demonstrate proficiency in Golang or Python to solve complex operational challenges and automate intricate workflows.
- Must have a deep understanding of container security best practices and apply them consistently across all deployments.
- Must have mandatory hands-on experience with Kubernetes, including proficiency with cert-manager or external-dns and GPU-enabled clusters.
- Must be experienced with infrastructure templating tools such as Helm or Kustomize for managing complex deployments.
- Must have production-grade experience with Infrastructure-as-Code tools like Terraform or CloudFormation.
- Must have hands-on experience with AWS services, including IAM, S3, EC2, and EKS, within live production products.
- Must possess experience with Google Cloud or Azure, with a significant advantage given for proven expertise in these platforms.
- Must have operational experience with PostgreSQL or similar relational database systems in production environments.
- Must be fluent with GitOps tooling such as Flux or Argo, as well as modern CI/CD systems like GitHub Actions.
- Must have a strong track record of troubleshooting complex infrastructure issues in high-availability environments.
- Must be comfortable working in a fast-paced, autonomous environment with clear accountability for outcomes.
Nice to have
- Direct, hands-on experience with GPU-enabled clusters is considered a bonus for this role.
Skills and tools
The role requires fluency with Kubernetes, Golang, Python, Helm, Kustomize, Terraform, CloudFormation, AWS, Google Cloud, Azure, PostgreSQL, Flux, Argo, and GitHub Actions.
Practical notes
This is a full-time engagement based in Germany with remote work options. Please verify all details on the official application page. All recruiting communications from Cresta originate exclusively from the @cresta.ai domain.