Principal CloudOps Engineer
Job description
About the role
Showpad is seeking a technical leader to manage and scale our cloud infrastructure. You will work hands-on with AWS environments to improve system reliability and support our growing platform capabilities. In this role, you will own the design and execution of cloud operations strategies that directly impact product delivery and customer experience. You will act as the primary technical authority for infrastructure decisions, driving initiatives that enhance performance, security, and cost optimization. This position requires a proactive mindset to identify risks before they escalate into critical outages or service degradation. You will mentor engineers across the platform by establishing best practices and standards for cloud operations. Your work will bridge the gap between development velocity and operational stability, ensuring our infrastructure can scale reliably with business growth. Through hands-on leadership, you will shape the future of our cloud-native ecosystem and define what operational excellence means for our organization.
Key facts
What you'll do
- Architect and implement advanced AWS-based infrastructure solutions to support exponential platform growth and evolving business requirements.
- Design and maintain infrastructure as code using CDK, CloudFormation, and related frameworks to ensure repeatability, version control, and auditability of all cloud resources.
- Lead complex incident response activities, performing deep root cause analysis and developing permanent remediation strategies for critical production issues.
- Reverse-engineer legacy and emergent systems to decipher functionality, dependencies, and data flows when documentation is incomplete or outdated.
- Assume responsibility for on-call rotations, providing timely responses to alerts and ensuring operational continuity with minimal business impact.
- Facilitate cross-functional collaboration to identify, prioritize, and remove technical blockers that impede product delivery or platform stability.
- Rapidly assimilate into our existing cloud environment, conducting thorough assessments to maintain and elevate performance standards from day one.
- Champion knowledge transfer sessions and create comprehensive documentation to elevate the technical capabilities of the entire engineering organization.
- Partner closely with development teams to embed operational considerations into the software development lifecycle from the initial design phase.
- Evaluate emerging technologies and determine their applicability for enhancing our infrastructure resilience, monitoring capabilities, and operational efficiency.
- Drive automation initiatives to reduce manual operational tasks, minimize human error, and free the team for higher-value engineering work.
- Establish and refine key performance indicators and operational benchmarks to measure infrastructure health, reliability, and cost-effectiveness.
- Contribute to the technical roadmap by providing insights on infrastructure trends, capacity planning, and long-term strategic alignment with business objectives.
- Serve as a subject matter expert during technical reviews, ensuring that infrastructure decisions align with security, compliance, and scalability requirements.
Requirements
- Utilize AI-native engineering tools such as Claude Code, OpenAI Codex, or Cursor on a daily basis to enhance productivity and solve complex technical challenges.
- Demonstrate extensive experience at a principal level managing and scaling highly complex AWS ecosystems in production environments.
- Provide evidence of successfully deploying infrastructure using AWS CDK or similar provisioning tools across multiple projects and environments.
- Exhibit strong proficiency in scripting and coding to automate operational tasks and build robust platform integrations with various services.
- Show a proven track record of troubleshooting and resolving issues within live production environments with minimal downtime and clear communication.
- Display the ability to work autonomously on ambiguous problems while rapidly adapting to new systems, technologies, and changing priorities.
- Maintain strong communication skills to articulate technical concepts to diverse audiences and actively own technical initiatives during critical escalations.
- Commit to working within the constraints of a hybrid model, including 2 days on-site in Ghent or Bucharest as required for team collaboration and strategic planning sessions.
Nice to have
- Hands-on experience operating or scaling AI/ML infrastructure and associated workload pipelines in production environments.
- Demonstrated background in high-throughput B2B or SaaS software environments with complex integrations and demanding performance requirements.
- History of modernizing mature, complex infrastructure landscapes while actively reducing technical debt and improving maintainability.
- Experience balancing deep individual technical contributions with effective mentorship of junior engineers and alignment with cross-functional stakeholders.
Practical notes
This role operates under a hybrid model requiring 2 days on-site presence in Ghent, Belgium or Bucharest, Romania to facilitate team collaboration and strategic planning. The position is full-time and requires availability to participate in on-call rotations as needed to ensure continuous system reliability. Candidates must be eligible to work in the designated location and comply with local regulations for employment.