Senior Software Engineer, Infrastructure
Job description
Senior Software Engineer, Infrastructure at Nectar Social.
About the role
We are looking for a Senior Software Engineer with a focus on infrastructure to take ownership of the reliability, scalability, and operational excellence of Nectar Social's production systems. This role involves managing high-volume data ingestion pipelines and real-time AI workloads on a platform experiencing rapid growth. As one of the first dedicated Site Reliability Engineers (SREs) in the company, you will have a significant impact on establishing the foundational reliability practices that support the company's scaling efforts. You will work closely with engineering teams to ensure systems are resilient, efficient, and capable of handling increasing demands, all while maintaining high standards for uptime and performance. Your work will directly influence the stability and efficiency of Nectar Social's platform, enabling the company to deliver a seamless experience to its users and partners.
Key facts
What you'll do
- Take ownership of the reliability and scalability of Nectar Social's production systems, which handle large volumes of social data and AI workloads in real-time.
- Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to measure system performance and reliability.
- Build and maintain observability tools, alerting systems, and on-call practices to ensure rapid detection and resolution of issues.
- Lead incident response efforts, coordinate blameless postmortems, and implement systemic improvements based on lessons learned to prevent recurrence of issues.
- Optimize system performance, improve cost efficiency, and plan capacity to support the platform's growth trajectory.
- Harden infrastructure-as-code, deployment pipelines, and CI/CD processes to ensure resilience, repeatability, and automation.
- Collaborate with engineering teams during system design to embed reliability principles from the outset, raising operational standards across the organization.
- Maintain and improve cloud infrastructure primarily on AWS, ensuring it meets security, compliance, and operational requirements.
- Automate operational tasks through scripting and tooling to reduce manual intervention and improve system consistency.
- Support security best practices in infrastructure management, including vulnerability mitigation and compliance adherence.
- Assist in onboarding new team members and sharing knowledge about reliability practices and infrastructure management.
- Develop capacity planning strategies to anticipate future infrastructure needs and avoid bottlenecks.
- Implement performance engineering techniques to ensure systems operate efficiently under load.
- Monitor system health continuously, analyzing metrics and logs to identify potential issues before they impact users.
- Work with cross-functional teams to improve deployment processes, reduce downtime, and increase deployment frequency.
- Contribute to the development of disaster recovery plans and business continuity strategies.
- Stay current with industry best practices and emerging technologies related to cloud infrastructure, automation, and reliability engineering.
Requirements
- 5+ years of experience operating production systems as an SRE, infrastructure engineer, or platform engineer.
- Proven experience in scaling databases, data infrastructure, or complex production platforms under significant load.
- Hands-on expertise with cloud infrastructure, especially AWS, and infrastructure-as-code tools such as Pulumi.
- Strong programming skills in scripting and automation languages to build tooling and operational services.
- Experience with monitoring, alerting, and observability tools to maintain system health and performance.
- Ability to operate effectively in fast-moving startup environments with high ownership and autonomy.
- A reliability-first mindset combined with pragmatism regarding velocity, costs, and operational trade-offs.
- Experience leading incident response efforts, conducting blameless postmortems, and implementing systemic improvements.
- Familiarity with database management, especially PostgreSQL, and data storage solutions like ClickHouse.
- Knowledge of capacity planning, performance engineering, and cost optimization at scale.
- Comfort working with infrastructure automation, deployment pipelines, and CI/CD practices.
- Strong communication skills to collaborate across teams and communicate reliability standards effectively.
- Ability to prioritize tasks and manage multiple projects simultaneously in a dynamic environment.
Nice to have
- Experience establishing or maturing an SRE practice at an early-stage or rapidly scaling company.
- Familiarity with the company's tech stack, including Pulumi, Postgres, ClickHouse, Turbopuffer, or Temporal.
- Background in capacity planning, performance engineering, or cost management at scale.
- Experience with security practices related to cloud infrastructure and data protection.
- Knowledge of disaster recovery planning and business continuity strategies.
- Prior experience working in a startup environment with high growth and evolving infrastructure needs.
Skills & tools
- AWS cloud platform for infrastructure hosting and management.
- Infrastructure-as-code tools, particularly Pulumi.
- PostgreSQL and other data storage solutions such as ClickHouse.
- Monitoring and observability tools for system health and alerting.
- CI/CD pipelines and automation scripting for deployment and operational tasks.
- Cloud security and compliance best practices.
- Incident management and postmortem analysis tools.
- Capacity planning and performance engineering techniques.
Practical notes
- This is an on-site role based in Palo Alto, CA.
- The company offers competitive compensation along with early equity options.
- Benefits include health, vision, and dental coverage, as well as a 401(k) match.
- The work schedule is hybrid, with four days in the office and three remote days per quarter after three months.
- The company provides comprehensive stipends, including a $1,000/month housing stipend for those living near the office, a $50/month mobile stipend, and a $50/month internet stipend for remote employees.
- Additional benefits include commuter benefits such as train stipends or parking permit reimbursement.
- Free lunch is available in the heart of University Ave. in Palo Alto, fostering a collaborative and engaging work environment.
About the company
We're living through a fundamental shift in how people discover, evaluate, and purchase products. The next generation doesn't respond to traditional marketing -- they build relationships with brands through authentic social interactions, seek recommendations from communities they trust, and expect personalized experiences that feel human, not corporate.