Site Reliability Engineer
asobbiUSA1w ago
EngineeringReliabilityremotecurated-jd
Job description
Site Reliability Engineer at asobbi.
About the role
A specialized AI infrastructure firm is establishing a new United States operations team following the acquisition of a major domestic client. We are seeking engineers to build out this capability from the start, focusing on automating large-scale compute infrastructure powered by renewable energy. This position emphasizes software engineering and automation to ensure platform reliability rather than AI model development.
Key facts
What you'll do
- Develop Python-based automation to handle incident response, routine operations, and runbook execution.
- Connect infrastructure APIs, observability platforms, and ITSM systems to improve workflow automation.
- Refine monitoring signals by implementing deduplication, correlation, and alert suppression.
- Create internal self-service tools, including CLI utilities, dashboards, and ChatOps integrations.
- Manage automation libraries and runbooks using version control.
- Translate findings from post-incident reviews into permanent tooling and operational standards.
Requirements
- Professional background in Platform Engineering, SRE, or production infrastructure management.
- Practical experience with monitoring and observability stacks like Grafana or Prometheus.
- History of managing on-call rotations and transforming manual procedures into automated tasks.
- Proficiency in Python for building integrations and automation scripts.
Nice to have
- Familiarity with GPU hardware, colocation facilities, or data center infrastructure.
- Experience integrating ITSM platforms such as Jira Service Management, Halo, or ServiceNow.
- Background in building ChatOps bots for Slack or Microsoft Teams.
- Knowledge of distributed tracing, logging, or OpenTelemetry.
- Experience with hypervisor control planes, IPAM, or DCIM integrations.
- Exposure to agent-based or LLM-assisted operational automation.
Skills & tools
- Python
- Prometheus
- Grafana
- Infrastructure APIs
- Version control
- Observability tooling
Practical notes
This role is fully remote but requires working within US time zones. Please apply directly to be considered for the position.