Site Reliability Engineer (SRE)
Job description
About the Role
Gradion is extending its global SRE capabilities for a long-term managed services client. This position places you within a follow-the-sun operational model, where platform stability and cloud infrastructure are paramount. Your responsibilities will span incident management, infrastructure orchestration, and automation. Success requires a high degree of self-direction, as you will undergo a structured onboarding process before operating independently across distributed time zones. You will be entrusted with maintaining the integrity of critical production systems, ensuring that services remain resilient and performant. This role demands a proactive mindset to identify potential issues before they escalate into major incidents. You will operate at the intersection of development and infrastructure, bridging the gap between code deployment and system reliability. Ultimately, your work will directly impact the customer experience for enterprise clients across the globe.
Our Commitment and Scale
Gradion operates as a strategic Digital Innovation and Deep Tech partner for enterprise clients worldwide. Our foundation is built on over 23 years of expertise, with a presence across three continents-Asia, Europe, and Africa. We maintain a team of more than 300 specialists in six countries, including Vietnam, Singapore, Thailand, Saudi Arabia, Germany, and Egypt. Our work serves over 100 enterprise customers, including notable unicorns such as Alaiko, HomeToGo, and Roadsurfer. We hold ISO 27001 certification and have been recognized as Vietnam's Best IT Company for eight consecutive years, ranking #1 for the past two years.
Key Facts
-
Location: Vietnam
-
Engagement: Long-term managed services
-
Compensation: 18,000,000 VND to 25,000,000 VND
What you'll do
- Orchestrate complex cloud infrastructure on AWS and GCP, automating resource lifecycles from provisioning through decommissioning.
- Engineer and maintain robust CI/CD pipelines and GitOps workflows to ensure rapid, reliable, and reversible deployments.
- Architect and operate observability platforms at production scale, tuning monitoring, logging, and alerting for signal over noise.
- Serve as a primary responder in the global follow-the-sun on-call rotation, driving incident resolution and continuity.
- Configure, deploy, and manage AI tooling and MCP servers in production environments, ensuring reliability and security.
- Develop internal automation tools and scripts using infrastructure-as-code principles to eliminate manual operational toil.
- Conduct in-depth post-incident reviews, extracting actionable insights and implementing preventative controls.
- Translate complex technical issues into clear narratives for both technical and non-technical stakeholders.
- Collaborate daily with engineering teams distributed across multiple time zones to align on operational goals.
- Drive continuous improvement initiatives to enhance platform reliability, performance, and efficiency.
Requirements
- Bring four or more years of hands-on experience in DevOps, SRE, or Platform Engineering within an international team setting.
- Demonstrate advanced mastery of Kubernetes, including cluster operations, troubleshooting, and intricate configuration management.
- Possess hands-on experience with AWS or GCP services for infrastructure provisioning, networking, and management.
- Show a clear understanding of networking fundamentals such as DNS, load balancing, firewalls, and VPC architectures.
- Exhibit proficiency in scripting and automation using Python or Bash to solve operational challenges.
- Have experience with CI/CD tools and GitOps-based delivery pipelines, ensuring code and infrastructure are versioned and tested.
- Prove competence in monitoring and observability systems, specifically Prometheus and ELK stack implementations.
- Communicate effectively in English to collaborate seamlessly with European stakeholders and documentation.
- Apply a self-directed and proactive approach to problem-solving, taking ownership of complex, ambiguous challenges.
- Adhere to strict operational standards and security best practices in all production activities.
Nice to have
- Experience configuring and managing MCP servers and AI tooling in production is valued.
- Backgrounds in AI enablement workflows or LLM infrastructure operations are considered advantageous.
- Familiarity with supporting eCommerce or SaaS platforms with high availability needs is beneficial.
- Experience working within the Frontastic or commercetools Frontend ecosystem is a plus.
Practical Notes
- This is a long-term managed services opportunity based in Ho Chi Minh City.
- The role requires participation in a global follow-the-sun on-call rotation.
- All application details can be found in the official process outlined in the What you'll do section.