Senior Infrastructure Engineer
Job description
About the role
Workato is seeking a Senior Infrastructure Engineer to join our Global Core Infrastructure team based in Singapore. You will collaborate with international teams to manage the cloud-native architecture that supports our automation platform, which handles billions of daily events for enterprise customers. In this capacity, you will play a pivotal role in ensuring the reliability, security, and scalability of the foundational systems that power our platform. The position requires a proactive approach to designing and maintaining the complex infrastructure that underpins our automation ecosystem. You will work closely with cross-functional partners to translate business requirements into robust technical solutions. This role is integral to the team responsible for driving the infrastructure that enables Workato's AI Lab initiatives. Your contributions will directly influence the performance and stability of the services that our enterprise customers rely on every day.
Key facts
What you'll do
- Architect and maintain a global, distributed cloud environment that ensures high availability and resilience.
- Coordinate with US-based engineering groups to execute platform initiatives and align on strategic technical roadmaps.
- Modify and refine open-source software to improve security, performance, and operational efficiency within our infrastructure stack.
- Create internal services by utilizing Linux networking primitives and system calls to optimize communication and data flow.
- Provision and oversee AWS resources using infrastructure as code tools like Terraform and Kubernetes to ensure consistency and scalability.
- Author and manage Kubernetes deployment configurations for core services, ensuring best practices are followed.
- Enhance CI/CD pipelines to ensure secure and efficient component delivery, reducing risk and accelerating release cycles.
- Conduct performance tuning on Linux systems and software to optimize resource utilization and application throughput.
- Develop internal tooling to automate operational workflows, reducing manual effort and potential for human error.
- Engage in incident response, system reliability analysis, and architecture reviews to continuously improve system robustness.
- Collaborate with development teams to provide infrastructure guidance and enable efficient application deployment.
- Monitor infrastructure health and implement proactive measures to prevent potential failures before they impact users.
- Evaluate and integrate new technologies to keep the infrastructure modern and aligned with industry standards.
- Document infrastructure designs and procedures to ensure knowledge sharing and operational continuity.
Requirements
- 7+ years of professional experience in systems or infrastructure engineering, demonstrating a proven track record in complex environments.
- Demonstrated background in cloud-native architectures and distributed systems, with a clear understanding of microservices principles.
- Deep knowledge of Linux internals, including system-level tuning, networking protocols, and process management.
- Practical experience managing AWS infrastructure at scale, including core services such as compute, storage, and networking.
- Proficiency in Kubernetes and Terraform, including the ability to write, debug, and optimize configurations for production workloads.
- Understanding of cloud security principles and network security, including best practices for securing infrastructure and data.
- Experience with CI/CD automation and deployment optimization, ensuring fast and reliable software delivery.
- Strong troubleshooting and analytical abilities to diagnose issues and implement effective solutions under pressure.
- Ability to work effectively in a fast-paced, collaborative environment with teams located in different time zones.
- Commitment to maintaining high standards of code quality, documentation, and operational excellence.
- Willingness to participate in on-call rotations to support critical infrastructure and respond to emergencies.
- Strong communication skills, both written and verbal, to articulate technical concepts to diverse stakeholders.
- A mindset focused on continuous improvement and automation to enhance operational efficiency.
Nice to have
- Proficiency in Golang for service development or tooling to build custom solutions and automation.
- Experience with data and messaging technologies such as Kafka, PostgreSQL, Redis, or ClickHouse to support diverse workloads.
- Familiarity with observability platforms like Prometheus, Grafana, or VictoriaMetrics for monitoring and troubleshooting.
- Knowledge of secrets management using tools like Vault to ensure secure handling of credentials.
- History of contributing to or maintaining open-source infrastructure projects, demonstrating technical leadership and community engagement.
Practical notes
- Reference ID: 2437.
- This role is based in Singapore and requires relocation or local candidacy.
- The engagement is for full-time employment, requiring standard working hours as determined by the team.
- Occasional travel may be required for team meetings or company events.
- Candidates must be eligible to work in Singapore without sponsorship requirements.
- The position is subject to background checks and standard hiring processes.
- All decisions regarding hiring and compensation are made in accordance with company policies and local regulations.