Senior Software Engineer, Release Infra
Job description
About the role
Join Brex as a Senior Software Engineer, Infrastructure (Release Engineering) and take ownership of the critical systems that orchestrate how our company ships software safely and at scale. You will design, build, and operate the release infrastructure that powers Brex's deployment pipelines, observability workflows, and incident management processes on a global scale. In this role, you will partner closely with product, platform, and operations teams to ensure every release meets our high standards for safety, speed, and reliability while our infrastructure scales securely to support tens of thousands of the world's best companies. You will drive the technical strategy for our release and observability systems, championing automation and reliability so our builders can move fast without sacrificing stability. You will proactively identify and mitigate risks across the release stack, define and monitor key performance indicators, and mentor engineers to elevate the entire organization's engineering craft. This is an opportunity to shape the future of how Brex delivers software and solidify your impact as a leader in infrastructure and reliability.
Key facts
What you'll do
Design, build, and maintain the release infrastructure that powers Brex's deployment pipelines and incident workflows.
Drive technical strategy and architecture for release and observability systems, making them more scalable, reliable, and secure.
Collaborate with product, engineering, and operations partners to ensure Brex's releases are safe, predictable, and low-friction.
Identify and deliver improvements to the end-to-end release process (from code merge to production) to reduce risk and cycle time.
Build and evolve tooling for observability and incident response, enabling fast detection, triage, and resolution.
Proactively identify and mitigate risks in our release and infrastructure stack, including performance, reliability, and security concerns.
Define, instrument, and monitor key metrics for release engineering (e.g., deployment frequency, change failure rate, MTTR) and use them to guide improvements.
Partner with other infrastructure and product teams to debug complex production issues and drive long-term fixes.
Contribute to and champion best practices in release engineering, reliability, and operational excellence across the organization.
Mentor other engineers on the team, providing technical guidance and code reviews to elevate the overall quality of our infrastructure.
Stay up-to-date on emerging tools and practices in release engineering, observability, and SRE, and bring relevant ideas into Brex's stack.
Evaluate and adopt new technologies that improve the reliability, security, and efficiency of our release workflows.
Own on-call responsibilities for release infrastructure and lead incident response efforts to restore service and prevent recurrence.
Work closely with security and compliance teams to ensure all release processes meet organizational and regulatory standards.
Requirements
7+ years of professional experience designing, building, and operating backend or infrastructure systems in production.
Strong proficiency in backend programming languages (e.g., Go, Java, Kotlin, or Python) with a focus on reliability and performance.
Hands-on experience with CI/CD and release pipelines (e.g., GitHub Actions, CircleCI, Buildkite, Argo, Spinnaker, Jenkins) including build, test, and deployment automation.
Experience architecting and operating scalable, highly available distributed systems that handle critical business workflows.
Deep understanding of infrastructure components such as containers, service meshes, and cloud platforms (e.g., AWS, GCP, or Azure).
Strong knowledge of observability tools and practices, including metrics, logs, and tracing (e.g., Prometheus, Grafana, Datadog, Sentry).
Experience with infrastructure as code and configuration management tools (e.g., Terraform, Ansible, Pulumi).
Solid understanding of security principles and practices relevant to release pipelines, artifact management, and access controls.
Nice to have
Experience with advanced deployment strategies such as canary releases, blue-green deployments, and feature flags.
Experience building internal developer platforms or self-service tooling for engineering teams.
Practical notes
This role is based in our San Francisco office and operates in a hybrid environment that combines the energy of the office with the flexibility of remote work. We require a minimum of three coordinated days in the office per week on Monday, Wednesday, and Thursday. As a perk, we also offer up to four weeks per year of fully remote work.