Senior AI Platform Reliability Engineer
HTC Global ServicesUSA5d ago
AIEngineeringPlatformReliabilityremotecurated-jd
Job description
Senior AI Platform Reliability Engineer at HTC Global Services.
About the role
We are seeking a Senior Site Reliability Engineer to manage and scale infrastructure specifically for enterprise AI and machine learning production environments. You will focus on the operational challenges of deploying AI services rather than model development, working in a hybrid capacity at one of our three office locations.
Key facts
What you'll do
- Architect and maintain highly available cloud and Kubernetes infrastructure for AI platforms.
- Manage the deployment of AI models, inference services, and agents into production.
- Build automated CI/CD pipelines using Harness or similar enterprise tools.
- Implement observability strategies including metrics, logs, and traces to monitor platform health.
- Lead capacity planning for compute, memory, and storage requirements for AI workloads.
- Troubleshoot complex distributed system issues involving databases, networking, and cloud services.
- Mentor engineering staff and define operational standards for reliability.
- Ensure all infrastructure meets enterprise security and compliance requirements.
Requirements
- 7+ years of experience in SRE, Platform Engineering, DevOps, or cloud infrastructure.
- Proven track record of operating AI, machine learning, or model-serving platforms at scale.
- Expert-level proficiency in Kubernetes administration and production operations.
- Strong technical background in Helm and Terraform.
- Hands-on experience with Google Cloud Platform, plus AWS or Azure.
- Proficiency in Python, Bash, and YAML for scripting and automation.
- Experience with production databases and messaging systems like PostgreSQL, Redis, Kafka, MongoDB, and Vault.
- Familiarity with observability tools such as OpenTelemetry, Prometheus, Grafana, Splunk, or AppDynamics.
- Ability to maintain 99.99% availability for platform solutions.
Nice to have
- Experience with generative AI, LLM inference, or AI agent platforms.
- Knowledge of AI-specific technologies like LiteLLM, Open WebUI, or vector databases.
- Background in operating GPU-enabled workloads or distributed inference services.
- Experience defining SLOs, SLIs, and error budgets for AI services.
- Familiarity with chaos engineering, progressive delivery, and automated resilience testing.
- Relevant cloud or Kubernetes certifications.
Skills & tools
- Kubernetes, Helm, Terraform, Google Cloud Platform, AWS, Azure
- Python, Bash, YAML
- Harness
- PostgreSQL, Redis, Kafka, MongoDB, Vault
- OpenTelemetry, Prometheus, Grafana, Splunk, AppDynamics
Practical notes
HTC Global Services provides a comprehensive benefits package including medical, dental, and vision insurance, 401(k) matching, paid time off, paid holidays, life and disability insurance, and professional development programs. The company is an equal opportunity employer committed to a diverse and inclusive workplace.