Site Reliability Engineer
kongItalyFull Time3d ago
PythonGolangAWSAzureGCPDockerKubernetesTerraformAnsibleJenkinsCI/CDAI
Job description
Site Reliability Engineer at kong.
About the role
The Site Reliability Engineering team at Kong builds and manages the large-scale infrastructure for our cloud services. This role focuses on ensuring high reliability and performance for critical customer applications. You will contribute to maintaining uptime and supporting product development teams.
Key facts
What you'll do
- Develop and maintain core infrastructure using Infrastructure as Code tools.
- Implement and manage monitoring, logging, and alerting systems to achieve high uptime.
- Troubleshoot and resolve production issues, then lead post-incident reviews to prevent future problems.
- Create automation to reduce manual tasks, improve system efficiency, and enable self-service for other engineering teams.
- Work with developers to integrate reliability and scalability practices throughout the application development process.
- Contribute to capacity planning, disaster recovery exercises, and security improvements.
- Participate in an on-call rotation to ensure platform availability.
Requirements
- Experience managing production systems on a major cloud platform (AWS, GCP, or Azure).
- Skilled in at least one programming or scripting language, such as Golang, Python, or Bash.
- Practical experience with containerization and orchestration technologies like Docker and Kubernetes.
- Understanding of Infrastructure as Code principles.
- Familiarity with CI/CD concepts and pipeline tools (e.g., GitLab CI, Jenkins).
- Knowledge of modern observability stacks (e.g., Prometheus, Grafana, ELK).
Nice to have
- Experience with Terraform.
Skills & tools
- AWS, GCP, Azure
- Golang, Python, Bash
- Docker, Kubernetes
- Terraform
- GitLab CI, Jenkins
- Prometheus, Grafana, ELK
Practical notes
Kong is building infrastructure for API and AI connectivity. The company's platform, Kong Konnect, helps organizations manage and secure intelligence flow across APIs and AI models.