Site Reliability Engineer
Job description
About the role
Plaud is seeking a Site Reliability Engineer to ensure the stability and performance of its AI products. This role involves leading efforts to improve system resilience and working with engineering teams to integrate reliability into product design. You will contribute to a company focused on amplifying human intelligence through innovative hardware and software. The position requires a proactive mindset dedicated to maintaining high availability and operational excellence. You will own the design of failure modes and the implementation of mitigation strategies before incidents occur. Collaboration with cross-functional partners will be central to defining and delivering robust infrastructure solutions. Your work will directly impact the reliability of products used by a global audience. This is an opportunity to shape the operational culture of a growing AI hardware and software company.
Key facts
Location: Singapore
Engagement: Full-time
(includes Employee Stock Ownership Plan, 401(k) with company match)
Team: Global Product R&D Center
What you'll do
- Architect and enforce reliability standards that govern the performance and scalability of Plaud.ai's infrastructure.
- Engineer and deploy automation scripts to eliminate manual toil and reduce the risk of operational errors in production environments.
- Optimize cloud resource allocation across AWS, GCP, and Azure to balance cost efficiency with system resilience.
- Construct and maintain observability pipelines that provide deep insights into system behavior and application performance.
- Coordinate with development teams to embed reliability principles into the software development lifecycle from the outset.
- Analyze historical incident data to identify patterns and implement preventative measures that reduce future risk.
- Oversee the configuration and management of Kubernetes clusters that host critical distributed applications.
- Lead the design of disaster recovery and business continuity plans to ensure service integrity during outages.
- Define and monitor Service Level Indicators to validate that systems meet established performance benchmarks.
- Mentor junior engineers on best practices for debugging complex issues in distributed systems.
- Evaluate and integrate new tooling that enhances system monitoring, logging, and tracing capabilities.
- Partner with product managers to align infrastructure roadmaps with evolving business objectives.
- Conduct post-incident reviews to extract actionable insights and drive changes across the technical organization.
- Ensure all infrastructure changes are delivered with version control and undergo rigorous testing before deployment.
Requirements
- Bring a minimum of 8 years of hands-on experience in Site Reliability Engineering, Infrastructure, or Platform Engineering roles.
- Demonstrate extensive operational knowledge of major public cloud providers including AWS, GCP, and Azure.
- Show practical expertise in deploying, managing, and scaling Kubernetes clusters in production environments.
- Prove a history of participating in on-call rotations and responding effectively to high-severity incidents.
- Exhibit strong proficiency in at least one modern programming language such as Go, Python, or Java.
- Display a solid understanding of networking fundamentals, load balancing, and security best practices in cloud environments.
- Illustrate the ability to work autonomously while communicating progress clearly to technical and non-technical stakeholders.
Nice to have
- Hold experience supporting AI/ML workloads or data-intensive platforms that require scalable infrastructure.
- Have a working familiarity with Service Level Objective and Service Level Agreement frameworks to measure reliability.
- Have contributed to fast-growing products where requirements evolve rapidly and adaptability is essential.
- Have managed systems distributed across multiple geographic regions with varying latency and compliance needs.
- Possess superior written and verbal communication skills for documenting processes and coordinating with global teams.
Practical notes
This role offers an Employee Stock Ownership Plan (ESOP) and a 401(k) retirement plan with company matching. Benefits include comprehensive medical, dental, and vision insurance, unlimited PTO, 13 paid holidays, and 12 weeks of fully paid parental leave. The workplace operates on a hybrid model, requiring a minimum of three in-office days per week. Employees receive access to AI tools like Cursor, GPT models, Gemini, and Claude, along with high-spec laptops and workstation setups. The position is based in Singapore and is a full-time engagement.