Senior Infrastructure Engineer/SRE
Job description
About the role
You will architect and operate the core infrastructure that powers Cresta's AI platform, enabling every conversation to become a competitive advantage for our customers. You will own the reliability, security, and scalability of the systems that allow engineering to move with speed and precision. You will collaborate closely with product and AI teams to translate business requirements into robust technical foundations. You will implement infrastructure-as-code patterns that standardize deployments across multiple cloud environments. You will design monitoring and observability solutions that provide deep insight into performance and security. You will automate repetitive operational tasks to free the organization to focus on high-impact innovation. You will champion best practices in container security and production resilience across distributed systems. You will help build the machine learning infrastructure that allows data scientists to train and deploy models at scale.
Key facts
What you'll do
Design and evolve developer toolchains that streamline engineering workflows and accelerate deployment pipelines across the organization.
Ensure unwavering reliability of multi-cloud Kubernetes clusters, pipelines, and underlying infrastructure through proactive maintenance and rapid response.
Define, implement, and monitor metrics, logging, analytics, and alerting frameworks that provide actionable insight into performance and security across all endpoints and applications.
Implement and govern infrastructure-as-code deployment tooling and supporting services that operate consistently across multiple cloud providers.
Drive automation of operations and engineering tasks to reduce manual toil and increase efficiency in routine processes.
Build and optimize machine learning infrastructure platforms that enable AI teams to train, test, and deploy models on large-scale datasets with reliability and speed.
Collaborate with security and compliance stakeholders to embed container security best practices into every layer of the infrastructure stack.
Lead incident response efforts and contribute to post-incident reviews that turn operational learnings into systemic improvements.
Partner closely with product and AI engineering teams to understand requirements and deliver infrastructure solutions that enable rapid experimentation and delivery.
Champion the adoption of GitOps tooling and CI/CD practices that promote traceability, repeatability, and high-velocity releases.
Own the design and operation of database systems, ensuring performance, durability, and scalability for critical workloads.
Establish standards for production readiness across cloud services, including AWS, Google Cloud, and Azure, and guide teams in consistent implementation.
Support the evolution of remote work tooling and infrastructure to maintain a high-trust, high-performance distributed engineering culture.
Contribute to technical documentation and knowledge sharing that improves onboarding and long-term maintainability of critical systems.
Requirements
5+ years of professional experience in DevOps, Site Reliability Engineering, Production Engineering, or a closely related field.
Deep proficiency with programming languages such as Golang or Python, with the ability to write clean, maintainable code for infrastructure automation.
Deep familiarity with container security best practices, including image scanning, runtime protection, and secure configuration.
Production experience working with Kubernetes and a deep understanding of the Kubernetes ecosystem, including popular open-source tooling such as cert-manager or external-dns.
Production experience with Kubernetes templating tools such as Helm or Kustomize for managing complex application deployments.
Production experience with infrastructure-as-code tools such as Terraform or CloudFormation for provisioning cloud resources.
Production experience working with AWS and core services such as IAM, S3, EC2, and EKS.
Production experience with additional cloud providers such as Google Cloud and Azure is highly valued.
Hands-on experience with database software such as PostgreSQL and understanding of database performance and reliability considerations.
Experience with GitOps tooling such as Flux or Argo for continuous delivery and operational control.
Experience with CI/CD platforms such as GitHub Actions to automate testing and deployment workflows.
Strong understanding of networking, load balancing, and security controls within cloud environments.
Ability to troubleshoot complex distributed systems and root-cause issues in production environments.
Commitment to writing tests, implementing monitoring, and maintaining documentation for infrastructure components.
Nice to have
Experience with GPU-enabled clusters for high-performance computing or AI workloads.
Experience contributing to open-source infrastructure projects or building internal platform tools.
Background in regulated industries where compliance and auditability are critical.
Practical notes
This is a full-time position based in the United States and eligible for remote work.
The role reports to the Infrastructure leadership team and operates across multi-cloud environments.
Candidates must be authorized to work in the United States without sponsorship at this time.
Application review will begin immediately and continue until the role is filled.