Senior Software Engineer, Site Reliability
Job description
About the role
At Upstart, we're united by a mission that matters: to radically reduce the cost and complexity of borrowing for all Americans. Every day, we bring creativity, experimentation, and advanced AI to reshape access to credit, helping millions move forward financially with clarity and confidence. As the leading AI lending marketplace, we partner with banks and credit unions to expand access to affordable credit through technology that's both radically intelligent and deeply human. Our platform runs over one million predictions per borrower using more than 1,800 signals, powering smarter, fairer decisions for millions of customers. We're proudly digital-first, giving most Upstarters the flexibility to do their best work from wherever they thrive, alongside teammates across 80+ cities in the US and Canada. Digital-first doesn't mean distant. We're intentional about in-person connection through team onsites, planning sessions, and moments that spark creativity and trust. And whether you choose to work primarily from home or collaborate in-person from one of our offices in Columbus, Austin, the Bay Area, or New York City (opening Summer 2026), you'll have the support to work in the way that works best for you. If you're energized by tackling meaningful problems, excited to innovate with purpose, and motivated by work that truly matters, we'd love to hear from you.
What you'll do
- Embody and share SRE principles at Upstart to establish a strong reliability culture across the organization.
- Exercise state-of-the-art SRE practices throughout the company to drive consistency and improvement in operational processes.
- Uphold a culture of visibility, ownership, and responsibility around service reliability for all systems, including customer-facing platforms.
- Implement standards for monitoring microservices, web apps, mobile apps, databases, Kubernetes clusters, and machine learning platforms in a fast-paced, high-growth environment.
- Improve incident response practices, both within the SRE team and across the entire organization, to reduce downtime and enhance resilience.
- Automate away toil that makes sense to be automated, focusing on sustainable and scalable solutions that reduce manual effort.
- Build and maintain internal tooling from scratch to support agile development techniques and operational needs, ensuring tools evolve with system complexity.
- Leverage strong software design and architecture skills to create robust, scalable systems that meet current and future demands.
- Work with multiple teams to deliver enterprise-wide solutions that meet strict reliability and performance standards, aligning with cross-functional objectives.
- Utilize fundamental knowledge of data structures and algorithms to solve complex operational challenges and optimize system behavior.
- Ensure proficiency with Infrastructure as Code using tools such as Terraform, CDK, and Cloudformation to manage cloud resources safely and efficiently.
- Maintain expertise in observability, monitoring, and reporting tools to provide actionable insights into system performance and customer experience.
Requirements
- Minimum of 6 years combined experience between Software Engineering, Site Reliability, and/or DevOps Engineering including CI/CD, TDD, internal tooling, observability, and other agile development practices.
- Proficiency coding in Python, Go, and JavaScript/TypeScript to develop and maintain critical systems that power our lending marketplace.
- Proficiency with Infrastructure as Code solutions including Terraform, CDK, and Cloudformation for managing cloud resources at scale.
- A software engineering background with experience building internal tooling from scratch and applying other agile development techniques in a regulated environment.
- Strong software design and architecture skills to create effective, maintainable, and resilient systems that support high availability.
- Fundamentally sound knowledge of data structures and algorithms to optimize performance, reliability, and scalability under load.
- Experience with on-call and incident management environments to handle production issues calmly and effectively.
- Experience with observability, monitoring, and reporting tools such as Datadog and Sumologic to ensure system health and surface insights quickly.
- Experience supporting SaaS software in a microservice-oriented cloud environment to maintain high availability and meet strict SLAs.
- Ability to work with multiple teams for enterprise-wide deliverables, collaborating closely with product, engineering, and operations to align on reliability goals.
- Must be authorized to work in the United States.
Practical notes
- Digital-first role with flexibility to work from wherever you thrive, alongside teammates across 80+ cities in the US and Canada.
- In-person connection is prioritized through team onsites, planning sessions, and moments that spark creativity and trust.
- Available work locations include Columbus, Austin, the Bay Area, and New York City (opening Summer 2026).