Senior Site Reliability Engineer
Job description
Senior Site Reliability Engineer at Nectar Social.
About the role
Nectar Social is developing an AI-native operating system designed to transform how brands and communities interact. We are seeking a Senior Site Reliability Engineer to manage the stability and performance of our production systems as we scale our social commerce platform. This role is responsible for ensuring that the underlying infrastructure supporting real-time AI agents and high-volume data streams remains robust, efficient, and secure. The ideal candidate will act as a guardian of operational excellence, balancing the demands of rapid feature development with the non-negotiable need for system reliability. You will own the design and execution of strategies that keep our platform available and performant under all conditions. This position requires a proactive mindset that anticipates failure modes before they impact customers and partners. You will be a key technical leader driving the standards and practices that define how Nectar Social delivers reliable experiences.
Key facts
What you'll do
- Oversee the reliability and scaling of production environments that process high-volume data streams and support autonomous AI agents in real time.
- Architect and enforce Service Level Objectives, Service Level Indicators, and error budgets to align technical performance with business goals.
- Engineer comprehensive observability frameworks that provide deep insight into system behavior and user impact.
- Lead the design and implementation of alerting mechanisms that ensure critical issues are surfaced promptly and accurately.
- Drive incident response operations, coordinating efforts to mitigate outages and minimize downtime for end users.
- Conduct in-depth postmortems that translate operational failures into actionable systemic improvements.
- Analyze and optimize cloud infrastructure to achieve peak performance while controlling costs and planning for future capacity needs.
- Refine infrastructure-as-code practices to ensure that environments are consistent, reproducible, and resilient across the lifecycle.
- Enhance CI/CD deployment pipelines to reduce risk and increase the frequency of safe, reliable releases into production.
- Partner closely with development teams to embed reliability standards and best practices directly into the software development lifecycle.
- Mentor engineers on operational principles and contribute to the growth of a strong site reliability culture.
- Evaluate and implement new technologies that improve the efficiency, scalability, and security of our infrastructure.
- Collaborate with cross-functional stakeholders to translate business requirements into robust technical specifications for infrastructure.
- Champion automation initiatives that eliminate manual toil and reduce the potential for human error in operations.
Requirements
- Bring 5 or more years of professional experience focused on Site Reliability Engineering, platform engineering, or infrastructure engineering in demanding production environments.
- Demonstrate a proven track record of successfully scaling complex data infrastructure and databases while maintaining performance under heavy load.
- Show practical experience with major cloud platforms, specifically deploying and managing services on AWS with mastery of core services.
- Exhibit proficiency in infrastructure-as-code tools, with a strong background using Pulumi to define and manage cloud resources.
- Display competence in programming to build automation, internal tools, and operational services that improve system reliability.
- Operate effectively with a high degree of autonomy in a fast-paced startup environment where priorities can shift quickly.
- Maintain a balanced approach that steadfastly prioritizes system stability without compromising the speed of development.
- Apply deep knowledge of databases and data architectures to ensure that storage solutions meet requirements for durability, consistency, and scale.
Nice to have
- Demonstrate prior experience building or maturing an SRE function within a startup or high-growth company.
- Show familiarity with the specific technology stack used at Nectar Social, including AWS, Pulumi, Postgres, ClickHouse, Turbopuffer, and Temporal.
- Bring expertise in performance engineering, detailed capacity planning, and large-scale cost optimization initiatives.
Skills & tools
- AWS
- Pulumi
- Postgres
- ClickHouse
- Turbopuffer
- Temporal
- Infrastructure-as-Code
- Observability and Alerting
- CI/CD Pipelines
Practical notes
- This is a hybrid role requiring 4 days per week in our Palo Alto office located in the heart of the city.
- Employees are granted 3 flex remote days per quarter after the completion of the first 3 months of employment.
- Comprehensive benefits include health, vision, and dental insurance to support your well-being and that of your family.
- The compensation package features 401(k) matching to help you build long-term financial security.
- A $1,000 monthly housing stipend is available for local residents to assist with the high cost of living in the Bay Area.
- Additional stipends include a $50 monthly mobile stipend and a $50 monthly internet stipend to support your work needs.
- Commuter support is provided through train stipends or parking permit reimbursements.
- Enjoy daily free lunch at our University Avenue office, fostering collaboration and community among the team.