
Senior Devops Engineer
Job description
About the role
Tomorrow.io is actively seeking a Senior DevOps Engineer to own the reliability, security, and operational efficiency of our global weather intelligence platform that powers critical decision-making for enterprise customers. In this position, you will architect and build self-service infrastructure platforms that abstract complexity from both weather scientists and software developers while ensuring robust performance and cost optimization. You will play a key role in integrating artificial intelligence directly into our day-to-day operational workflows to accelerate development cycles and improve system observability. The role requires you to manage and maintain sophisticated cloud-native Kubernetes environments in parallel with high-performance scientific computing clusters to meet the demands of precise atmospheric modeling. You will work closely with cross-functional teams including spacecraft mission operations personnel to ensure that our infrastructure meets the stringent requirements of real-time data ingestion and processing. This position involves designing adaptive infrastructure that can scale rapidly as we expand our weather data coverage and analytics capabilities globally. You will be responsible for establishing best practices that enable our engineering and scientific teams to operate autonomously and deploy changes safely and frequently.
Key facts
What you'll do
- Develop and implement AI-driven tools that enhance development velocity and streamline operations workflows across the engineering organization.
- Partner with interdisciplinary teams including weather scientists, software engineers, and Spacecraft Mission Operations to analyze and optimize service performance while managing infrastructure costs.
- Design and maintain adaptive cloud infrastructure architectures that support aggressive business growth and evolving scientific computing requirements.
- Create intuitive self-service platforms and internal developer portals that allow developers and research scientists to provision resources and manage workloads independently.
- Manage and optimize scientific computing workloads on High-Performance Computing clusters using SLURM workload manager in conjunction with containerized Kubernetes environments.
- Integrate MLOps practices and patterns to enable the reliable deployment of GPU-intensive machine learning models within Kubernetes clusters.
- Participate in rotational on-call duties to ensure continuous production availability and rapid response to critical infrastructure incidents.
- Implement infrastructure monitoring and alerting strategies that provide actionable insights into system performance and resource utilization across hybrid cloud environments.
- Automate routine operational tasks and infrastructure management processes to reduce manual overhead and improve team efficiency.
- Collaborate with security and compliance stakeholders to ensure that infrastructure implementations adhere to organizational policies and regulatory requirements.
- Evaluate and recommend new tools and technologies that can enhance the capabilities of our cloud and HPC infrastructure stack.
- Contribute to the development of technical documentation and runbooks to support operational procedures and knowledge transfer.
- Mentor junior engineers and DevOps practitioners by sharing expertise in cloud technologies, infrastructure patterns, and automation strategies.
- Drive initiatives to improve system resilience, scalability, and disaster recovery capabilities across all production environments.
Requirements
- Possess a minimum of 6 years of professional experience as a Platform Engineer, DevOps Engineer, or Site Reliability Engineer within containerized cloud environments.
- Demonstrate proficiency with major cloud platforms such as AWS, Google Cloud Platform, or Microsoft Azure and Infrastructure as Code tools including Terraform or Crossplane.
- Bring experience working within fast-growing, cloud-native startups or scale-up companies where agility and rapid iteration are essential.
- Utilize AI coding agents such as Claude Code or GitHub Copilot on a daily basis to enhance personal productivity and code quality.
- Apply CI/CD methodologies and implement Kubernetes deployment strategies including blue-green deployments, canary releases, and rolling updates.
- Write clean, maintainable code and demonstrate proficiency in programming languages such as Python, Node.js, and Go.
- Configure and manage monitoring systems and observability platforms including Datadog, Prometheus, Grafana, or the ELK Stack to gain insight into system health.
- Exhibit strong collaboration skills to work effectively across distributed teams and engage with research and development stakeholders.
- This position is restricted to United States citizens, permanent residents, or protected individuals due to United States export control laws and regulatory requirements.
Nice to have
- Hands-on experience building agentic DevOps workflows that enable autonomous system operations and self-healing capabilities.
- Background in High-Performance Computing environments including practical experience with SLURM, AWS ParallelCluster, or Azure CycleCloud for managing compute clusters.
- Familiarity with parallel file systems such as Lustre or Network File System implementations used in scientific computing scenarios.
Practical notes
- The compensation package for this role includes comprehensive health benefits and unlimited paid time off to support work-life balance.
- Tomorrow.io maintains an Equal Employment Opportunity and Affirmative Action employment policy and actively participates in the E-Verify program for employment eligibility verification.
- Reasonable accommodations are available for candidates with disabilities by contacting the dedicated email address jobs@tomorrow.io during the application process.
- This position requires the successful candidate to be physically located within the Washington, District of Columbia, United States region due to the nature of the work and client requirements.
- The role involves rotational participation in on-call schedules to ensure continuous coverage and rapid response to production incidents across all supported infrastructure platforms.
- Candidates should be prepared to engage with United States-based security and compliance frameworks due to the export control restrictions governing this position.
- Successful candidates will undergo a standard employment verification process as part of the final hiring considerations for this role.
- The responsibilities of this role may evolve over time to address changing business needs, technological advancements, and organizational priorities in the weather intelligence sector.