Middle - Senior DevOps Engineer
Job description
About the role
Gradion positions you as a key operator within a global, follow-the-skin SRE function that underpins a strategic client engagement extending through 2028. You will act as the primary guardian of platform availability, ensuring that critical systems remain stable and performant across demanding European and global time zones. This role requires a self-directed engineer who thrives in a fast-moving, internationally distributed environment and takes ownership of complex operational challenges. You will work closely with an internal SRE team during a structured onboarding period to align with practices and standards before operating with full independence. Your daily work will involve managing cloud infrastructure, automating operations, and responding to incidents with precision and accountability. You will contribute directly to the reliability and security posture of large-scale enterprise platforms that drive digital innovation. This is an opportunity to deepen your expertise in production operations while shaping the long-term resilience of a high-impact client platform.
Key facts
What you'll do
Design and maintain robust platform availability by monitoring, triaging, and resolving incidents within strict SLA windows across global time zones.
Manage the full lifecycle of cloud infrastructure on AWS and/or GCP, including provisioning, scaling, configuration, and day-to-day operational health.
Build, refine, and optimize CI/CD pipelines and GitOps workflows to accelerate delivery while maintaining stability and compliance.
Operate observability systems at production scale, tuning monitoring, logging, and alerting to balance sensitivity and noise.
Participate in a 24/7 on-call rotation as part of the follow-the-sun coverage model, ensuring timely response and coordination across regions.
Configure, deploy, and manage AI tooling and MCP servers in production environments with attention to reliability and security.
Lead infrastructure automation initiatives, writing scripts and building internal tools that simplify operations for distributed teams.
Author clear, factual post-incident reviews and synthesize findings into the monthly operational report used by leadership and stakeholders.
Collaborate daily with engineering teams across multiple time zones, aligning on priorities, runbooks, and improvement actions.
Champion best practices for security, cost governance, and performance within cloud platforms and operational processes.
Requirements
Bring a minimum of 4 years of hands-on experience in DevOps, SRE, or Platform Engineering within an international team setting.
Demonstrate solid Kubernetes knowledge, including cluster operations, troubleshooting, and configuration in diverse production scenarios.
Provide evidence of hands-on cloud experience with AWS and/or GCP, covering core services and operational best practices.
Show a strong grasp of networking fundamentals such as DNS, load balancing, firewalls, and VPC design principles.
Exhibit scripting and automation proficiency in languages such as Python or Bash to solve operational problems efficiently.
Have direct experience with CI/CD tools and GitOps-based delivery pipelines that support rapid, reliable deployments.
Possess a working knowledge of monitoring and observability systems like Prometheus, ELK, or equivalent platforms.
Communicate effectively in English on a daily basis, as collaboration with European stakeholders is a core requirement of the role.
Operate as a self-directed and proactive professional who asks the right questions and drives issues to resolution without waiting for direction.
Maintain a strong sense of ownership for operational outcomes, reliability targets, and continuous improvement in production environments.
Nice to have
Bring experience configuring and managing MCP servers and AI tooling in production environments with measurable reliability.
Show background in AI enablement workflows or LLM infrastructure operations within cloud platforms.
Have a background supporting eCommerce or SaaS platforms with complex integration and uptime requirements.
Demonstrate familiarity with the Frontastic or commercetools frontend ecosystem and related integration patterns.
Practical notes
The engagement is full-time based from Ho Chi Minh City.
This role requires participation in a global follow-the-sun on-call rotation.
No visa sponsorship details are provided in this listing.
There are no published deadlines for submission specified in the source.