Senior SRE, Software Engineering
Job description
About the role
You will own the reliability and scalability of the core systems that power our consumer finance operations on a daily basis. This role requires you to step into production incidents with calm authority and drive solutions that prevent recurrence. You will define what reliability means for our organization and codify it into processes and tools that the entire engineering team uses. Your technical decisions will shape how we scale from thousands to millions of borrowers without sacrificing stability. You will be the architect of our on-call culture and the designer of how we respond when things go wrong. Ultimately, you will ensure that the product remains available and performant for the millions of consumers who depend on us.
Key facts
What you'll do
Lead incident response and establish sustainable on-call practices, including comprehensive runbooks, blameless postmortems, and systematic improvements that reduce MTTR.
Develop and maintain self-service observability solutions using modern monitoring tools that provide actionable insights for troubleshooting and performance optimization.
Create and maintain infrastructure as code (using Terraform, CloudFormation) that allows for consistent, scalable, and secure cloud environments on AWS.
Partner closely with feature teams to architect resilient infrastructure for critical components (databases, networking, async workflows, data pipelines) that scale seamlessly.
Work closely with DevX to design and implement robust CI/CD pipelines with advanced deployment strategies (blue/green, canary) that enable teams to ship confidently and rapidly.
Advocate for best practices early in feature design, ensuring we design with reliability in mind and future-proof our services.
Implement advanced database optimization techniques to handle the scale and complexity of our growing data layer.
Design and manage async workflow infrastructure to ensure reliable processing of high-volume operations across distributed systems.
Own the reliability of data pipelines, ensuring data integrity, timely processing, and availability for downstream consumers and decision systems.
Establish clear service level objectives and service level indicators, aligning technical metrics with business outcomes.
Build and maintain comprehensive documentation for systems and processes to enable efficient onboarding and knowledge transfer.
Collaborate with security and compliance teams to ensure infrastructure and processes meet regulatory and organizational standards.
Mentor junior engineers and SREs, fostering a culture of learning and shared ownership of system reliability.
Continuously evaluate new tools and technologies to improve the efficiency, resilience, and cost-effectiveness of our infrastructure.
Requirements
Expertise leading incident response for high-availability production systems, thorough root cause analysis, and fostering blameless postmortem culture.
Experience designing highly available deployment architectures across multiple targets (e.g. EC2, Fargate), with expertise in auto-scaling, health checks, and graceful degradation strategies.
Strong knowledge of AWS cloud services and infrastructure-as-code practices using tools like Terraform and CloudFormation.
Track record of implementing effective monitoring & observability solutions (e.g. Datadog, Prometheus, ELK), and evangelizing best practices.
Experience with CI/CD pipelines and automation to enable reliable, efficient deployments.
Excellent communication skills with experience documenting processes and collaborating across engineering teams.
Bachelor's degree in Computer Science, Engineering, or a related technical field or equivalent practical experience.
5+ years of experience in software engineering, SRE, or platform engineering roles with a focus on production systems.
Nice to have
Experience with database optimization and tuning in high-throughput environments.
Background in consumer finance or experience working with regulated industries.
Familiarity with AI-assisted development tools and how they integrate into engineering workflows.
Practical notes
Location based in New York City.
Full-time engagement.
Office in Nolita with a mandate to work in person at least three days a week.
About January
January is hiring for Senior SRE, Software Engineering. The listing location is New York City.
This Senior SRE, Software Engineering opening is posted for New York City.