Site Reliability Engineer
Job description
.
About the role
Alpaca operates as a US-headquartered global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, and 24/5 trading. The company provides licensed financial services to hundreds of financial institutions across 40 countries through its institutional-grade APIs. This infrastructure supports broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, managing over 10 million brokerage accounts. Alpaca employs a globally distributed team of engineers, traders, and brokerage professionals dedicated to opening financial services for everyone. The company maintains a commitment to open-source contributions and community growth while enhancing its developer-friendly API and robust infrastructure. Alpaca is backed by $400 million in funding from investors including Portage Ventures, Spark Capital, Tribe Capital, Social use, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.
The Site Reliability Engineer role focuses on end-to-end reliability for trading infrastructure. You will shape intake, build, review, and ship phases to protect broker-dealer services for global markets. The position translates complex constraints into clear guardrails, enabling product and infrastructure teams to move safely. You will ensure stability across critical paths while mentoring peers on database and cloud practices within a globally distributed squad. This role reports as a remote position based in EMEA. It is a full-time engagement with annual compensation ranging from 136,000 USD to 176,000 USD.
You will design intake pipelines that capture system metrics and events, enabling fast detection before issues affect customers. You will champion build standards for cloud resources and Kubernetes objects, enforcing GitOps workflows that support broker-dealer compliance. You will drive review routines for observability, defining SLIs and error budgets that align product decisions with reliability targets. You will orchestrate shipping strategies for messaging and data layers, ensuring changes roll out safely across global regions. You will guard PostgreSQL performance and availability, handling schema reviews, online migrations, and CDC pipelines on critical paths. You will elevate partner reliability through collaboration, guiding other teams on debugging, incident practices, and capacity planning. You will mentor engineers by pairing on code reviews and design sessions, strengthening database and SRE fundamentals across the group. You will expand observability coverage, correlating metrics, logs, traces, and alerts to clarify behavior in complex distributed flows.
Key facts
What you'll do
- Shape intake pipelines that capture system metrics and events, enabling fast detection before issues affect customers.
- Champion build standards for cloud resources and Kubernetes objects, enforcing GitOps workflows that support broker-dealer compliance.
- Drive review routines for observability, defining SLIs and error budgets that align product decisions with reliability targets.
- Orchestrate shipping strategies for messaging and data layers, ensuring changes roll out safely across global regions.
- Guard PostgreSQL performance and availability, handling schema reviews, online migrations, and CDC pipelines on critical paths.
- Elevate partner reliability through collaboration, guiding other teams on debugging, incident practices, and capacity planning.
- Mentor engineers by pairing on code reviews and design sessions, strengthening database and SRE fundamentals across the group.
- Expand observability coverage, correlating metrics, logs, traces, and alerts to clarify behavior in complex distributed flows.
- Translate complex constraints into clear guardrails that enable product and infrastructure teams to move safely.
- Ensure stability across critical paths while maintaining alignment with broker-dealer regulatory expectations.
- Implement structured incident response practices, including calm debugging and actionable postmortems.
- Collaborate closely with product and infrastructure teams to evolve reliability practices in a fast-moving environment.
Requirements
- Bring 4+ years of experience operating production systems, with clear ownership of incidents, postmortems, and follow-through that closes reliability loops.
- Operate Kubernetes clusters in production, using infrastructure as code and GitOps pipelines to manage cloud resources safely.
- Apply solid PostgreSQL expertise, including query plans, pg_stat views, indexing trade-offs, and safe online migrations at scale.
- Understand cloud networking fundamentals, including VPCs, routing, load balancing, DNS, and TLS, with skill in debugging cross-service issues.
- Work confidently in a modern observability stack and operate efficiently at the Linux command line during high-pressure incidents.
- Practice incident response with structured debugging, writing postmortems that generate concrete reliability improvements.
- Write Go or Python code with strong written and verbal communication to coordinate with global product and infrastructure teams.
- Meet the bar Nice to have
- Experience contributing to open source projects that touch financial services infrastructure and broker-dealer workflows.
- Skills and tools including Kubernetes, PostgreSQL, Git, Go, Python, Prometheus, Grafana, Loki, Alertmanager, Kafka, and Snowflake.
Practical notes
This is a remote role based in EMEA. The position is full-time with an annual compensation range of 136,000 USD to 176,000 USD. The role reports into EMEA remote operations. Applicants should expect to align with working hours common to EMEA-based teams. No additional travel or visa requirements are specified in this listing. Please refer to the official apply page for further practical details, deadlines, and application instructions.