Staff Site Reliability Engineer
Job description
About the role
You will own the design and execution of infrastructure that underpins critical insurance claims workflows, ensuring the platform remains resilient and performant for large-scale enterprise users. You will architect and automate the provisioning of environments so that product teams can move quickly without sacrificing stability or security. You will implement robust monitoring and observability practices that provide clarity into complex, distributed systems. You will partner closely with cross-functional teams to embed security and compliance controls directly into the delivery pipeline. You will lead incident response efforts and on call rotations to maintain high service standards. You will mentor engineers across the stack to elevate operational excellence across the organization. You will help define the long-term strategy for our SaaS infrastructure as we scale to support billions of dollars in claims transactions.
Key facts
What you'll do
Orchestrate the provisioning and scaling of infrastructure required to deliver the Assured platform, evolving existing products and enabling new initiatives that respond to market demands.
Automate the configuration and maintenance of our platform and services, building tools that emphasize consistency, repeatability, and idempotency across environments.
Design and implement sustainable methods for monitoring, managing, and scaling our platforms and services to meet growing customer and regulatory expectations.
Collaborate with product and engineering teams to improve observability across different product areas, ensuring that insights drive better decision-making and faster resolution.
Identify, recommend, and implement strategies and controls to comply with security regulations and industry best-practices, embedding them into infrastructure as code.
Provide incident support and participate in on call rotation, responding promptly to production issues and driving remediation efforts to minimize customer impact.
Lead, mentor, and coach other engineers, fostering a culture of learning, knowledge sharing, and ownership across the growing engineering organization.
Evaluate and adopt technologies and patterns that enhance reliability and scalability, balancing innovation with pragmatic delivery and operational simplicity.
Partner with development teams to ensure that platform changes are well-understood, documented, and rolled out with appropriate safeguards and rollback strategies.
Champion automation where it adds value, while also recognizing scenarios where manual oversight or alternative approaches are more appropriate.
Take initiative to identify and resolve complex operational problems, coordinating across teams to address cross-cutting technical risks and dependencies.
Promote engineering excellence through thorough code reviews, clear documentation, technical guidance, and active mentorship of peers and junior engineers.
Persistently pursue resolution of difficult operational roadblocks, efficiently navigating constraints and enlisting support from other teams when necessary.
Work with data platforms and database technologies to ensure that database solutions are highly available, scalable, and aligned with business and compliance requirements, with a focus on PostgreSQL.
Requirements
You bring a strong engineering background and a history of designing systems that are reliable, scalable, and maintainable in production environments.
You have experience working in a start-up environment where ambiguity is common and you thrive by taking ownership of ambiguous problems.
You are comfortable autonomously building and scaling high-quality products and environments from early stages through to maturity and steady state.
You have a demonstrated ability to automate processes, but you can also discern when automation is not the best solution and present thoughtful alternatives.
You take the initiative to identify and solve important problems, coordinating with others on cross-cutting technical and operational challenges.
You are self-motivated and require minimal direction or oversight, while still collaborating effectively with distributed teams.
You have experience designing, implementing, and maintaining highly available and scalable database solutions, ideally with PostgreSQL, and understand the tradeoffs involved.
You have a passion for implementing and supporting observability platforms, defining and tracking SLOs, and building effective monitoring and alerting systems.
You have experience working in scaling environments with Terraform, AWS, and Kubernetes, and you understand the implications of infrastructure decisions on security, cost, and reliability.
Nice to have
Experience contributing to open source projects that demonstrate collaboration and technical rigor.
Familiarity with insurance domain concepts or prior work in financial services or regulated environments.
Hands-on experience with security and compliance frameworks relevant to the insurance industry.
Practical notes
This is a full-time remote position.
Our interviews are conducted through verified company channels, and we only contact candidates from official @assured.claims email addresses. If you are unsure about the legitimacy of a message, please contact recruiting-ops@assured.claims before sharing any personal information.