Site Reliability Engineer
outpostRemoteContract2d ago
TypeScriptPythonReactNext.jsNode.jsExpressGCPDockerTerraformCI/CDPostgreSQLComputer Vision
Job description
Site Reliability Engineer at outpost
About the role
Outpost is building the core infrastructure for freight operations, utilizing a blend of real estate, operations, and technology. We are seeking a Site Reliability Engineer to ensure the stability and performance of our mission-critical systems that keep freight moving. This role is essential for scaling our platform and maintaining customer trust.
Key facts
What you'll do
- Take ownership of reliability objectives for our backend services, applications, and computer vision pipeline, focusing on metrics like Mean Time To Detect, Mean Time To Mitigate, and Mean Time To Recover.
- Enhance our monitoring and alerting systems, developing automated remediation capabilities to manage on-call responsibilities efficiently.
- Collaborate with engineering teams to integrate AI agents for alert triaging and routine issue resolution.
- Optimize and secure our Google Cloud Platform (GCP) infrastructure, including Cloud Run, Cloud SQL, and GCS, for cost and performance as usage grows.
- Manage database scalability and performance, addressing connection pooling, query optimization, indexing, and capacity planning for PostgreSQL.
- Improve the reliability of our machine learning training and monitoring infrastructure.
- Conduct blameless postmortems to identify and address root causes of incidents.
- Participate in an on-call rotation.
Requirements
- A minimum of 4 years of experience in a Site Reliability Engineering, infrastructure, or backend engineering capacity, including production on-call responsibilities.
- Proficient with a major cloud provider, with a preference for GCP, including experience with compute services, managed databases, object storage, and networking.
- Proven experience building and managing monitoring, alerting, and observability stacks (e.g., Grafana, Prometheus, Datadog).
- Strong scripting and automation skills in languages such as Python or Bash.
- Familiarity with containerized applications using Docker and CI/CD pipelines.
- Demonstrated success in reducing incident frequency or improving reliability metrics.
- Excellent communication skills, capable of working effectively with both technical and non-technical colleagues and escalating issues appropriately.
- Clear and effective written and verbal communication in English, with strong asynchronous communication abilities.
Nice to have
- Experience with machine learning or data infrastructure, including training pipelines, model monitoring, and feature stores.
- Experience developing or integrating AI agents for operational automation.
- Familiarity with Infrastructure-as-Code tools like Terraform.
- Expertise in PostgreSQL performance tuning at scale.
- Background supporting physical or IoT systems.
- Experience with bare-metal infrastructure in colocation environments.
Skills & tools
- GCP (Cloud Run, Cloud SQL, GCS)
- PostgreSQL
- Docker
- Grafana, Prometheus, Datadog (or similar)
- Python, Bash
- TypeScript, Node.js, Next.js, React, Apollo Server, Express (familiarity with the stack is beneficial)
Practical notes
- This is a full-time contract position.
- Employment will be managed through an agency.
- Outpost is an Equal Opportunity Employer.