Infrastructure Engineer/SRE
Job description
About the role
You will design, build, and operate the core infrastructure that powers Cresta's AI platform and enables rapid, secure product execution. You will own the reliability, security, and scalability of multi-cloud Kubernetes environments and the developer toolchain that empowers engineering workflows. You will implement infrastructure-as-code, automate operations, and build machine learning infrastructure to support large-scale AI training and deployment. You will establish robust metrics, logging, and alerting to ensure performance and security across all endpoints. You will collaborate closely with engineers while maintaining high autonomy and clear ownership of defined responsibilities. You will help define and evolve best practices that allow the organization to move quickly while keeping systems resilient.
Key facts
What you'll do
- Partner with engineering teams to develop and maintain developer toolchains that streamline workflows and improve deployment infrastructure.
- Ensure high reliability, availability, and performance of multi-cloud Kubernetes clusters and associated CI/CD pipelines.
- Design, implement, and refine metrics, logging, analytics, and alerting systems for performance and security across all endpoints and applications.
- Build and manage infrastructure-as-code deployments supporting multiple cloud providers and their native services.
- Drive automation of operational and engineering tasks to reduce manual effort and increase efficiency.
- Build and optimize machine learning infrastructure that enables AI teams to train, test, and deploy models on large-scale datasets.
- Implement and evolve GitOps practices using tools such as Flux or Argo to manage declarative configuration and deployment.
- Maintain and secure containerized environments following container security best practices and compliance standards.
- Troubleshoot complex issues in production systems using observability data and collaborate on root cause analysis.
- Support the adoption of Kubernetes templating tools such as Helm and Kustomize across services and teams.
- Work with database systems such as PostgreSQL to ensure performance, reliability, and scalability for application needs.
- Contribute to the development and maintenance of internal platforms that enable faster feature delivery and operational stability.
Requirements
- Bring 5+ years of professional experience in DevOps, Site Reliability Engineering, Production Engineering, or a closely related field.
- Demonstrate deep proficiency with programming languages such as Golang or Python for automation and tooling.
- Show deep familiarity with container security best practices and secure deployment patterns for containerized workloads.
- Provide production experience with Kubernetes and a strong understanding of the Kubernetes ecosystem, including associated open-source tools.
- Demonstrate hands-on experience with Kubernetes templating tools such as Helm or Kustomize in production environments.
- Show production experience with infrastructure-as-code tools such as Terraform or CloudFormation for provisioning cloud resources.
- Provide production experience managing applications and services on AWS, including IAM, S3, EC2, and EKS.
- Demonstrate experience with additional cloud providers such as Google Cloud and Azure as a bonus.
- Bring experience with production database systems such as PostgreSQL.
- Show experience with GitOps tooling and operational patterns using systems such as Flux or Argo.
- Demonstrate experience with CI/CD pipelines and tooling such as GitHub Actions.
- Provide evidence of working effectively in a collaborative, autonomous team environment with clear ownership and accountability.
Nice to have
- Experience with GPU-enabled Kubernetes clusters and related scheduling and orchestration tooling.
- Familiarity with Cresta's AI and machine learning infrastructure priorities and related open-source projects.
- Contributions to or knowledge of large-scale distributed systems in production environments.
- Experience with monitoring and observability platforms such as Prometheus, Grafana, or similar tools.
- Background in customer experience or AI-driven applications in contact center or conversational AI domains.
Practical notes
- This role is based in Taiwan and is offered as a remote position.
- Compensation for this role includes a base salary, equity, and a variety of benefits. Actual base salaries will be based on candidate-specific factors, including experience, skillset, and location, and local minimum pay requirements as applicable.
- We are actively hiring for this role and encourage qualified candidates to apply promptly.
- Only candidates who match the outlined requirements and demonstrate relevant production experience will be considered for further review.