Senior Devops Engineer
Tomorrow.ioUSA1w ago
Engineeringremotecurated-jd
Job description
Senior Devops Engineer at Tomorrow.io.
About the role
Tomorrow.io is looking for a Senior DevOps Engineer to manage the reliability, security, and efficiency of our global weather intelligence platform. You will build self-service infrastructure, integrate AI into operational workflows, and support both cloud-native Kubernetes environments and high-performance scientific computing clusters.
Key facts
What you'll do
- Develop and implement AI-driven tools to improve development and operations efficiency.
- Partner with weather scientists, software engineers, and Spacecraft Mission Operations to optimize service performance and cost.
- Design and maintain adaptive cloud infrastructure to support rapid business growth.
- Create self-service platforms that allow developers and scientists to work autonomously.
- Manage scientific computing workloads on HPC clusters using SLURM alongside Kubernetes.
- Integrate MLOps practices for deploying GPU-based models on Kubernetes.
- Participate in on-call rotations to maintain production availability.
Requirements
- Minimum 6 years of experience as a Platform, DevOps, or SRE engineer in a containerized cloud environment.
- Proficiency with AWS, GCP, or Azure and Infrastructure as Code tools like Terraform or Crossplane.
- Experience working within fast-growing, cloud-native startups or scale-up companies.
- Daily, hands-on experience using AI coding agents like Claude Code or Copilot.
- Experience with CI/CD methodologies and Kubernetes deployment strategies.
- Proficiency in Python, Node.js, and Go.
- Experience configuring monitoring systems such as Datadog, Prometheus, Grafana, or the ELK Stack.
- Ability to collaborate effectively across distributed teams and R&D stakeholders.
- This position is restricted to U.S. citizens, permanent residents, or protected individuals due to U.S. export control laws.
Nice to have
- Experience building agentic DevOps workflows.
- Background in HPC or scientific computing environments including Slurm, AWS ParallelCluster, or Azure CycleCloud.
- Familiarity with parallel filesystems like Lustre or NFS.
Skills & tools
- Cloud: AWS, GCP, Azure
- IaC: Terraform, Crossplane
- Orchestration: Kubernetes
- Languages: Python, Node.js, Go
- Monitoring: Datadog, Prometheus, Grafana, ELK Stack
- HPC: SLURM, ParallelCluster, CycleCloud, Lustre, NFS
- AI: AI coding agents, MLOps
Practical notes
- Compensation includes comprehensive health benefits and unlimited paid time off.
- Tomorrow.io is an Equal Employment Opportunity and Affirmative Action employer and participates in E-Verify.
- Reasonable accommodations are available for candidates with disabilities by contacting jobs@tomorrow.io.