
Systems Engineer
Job description
About the role
Tomorrow.io is a weather intelligence company that builds platforms for forecasting and environmental data analysis. The Systems Engineer will support the infrastructure that processes and delivers weather data to customers worldwide. This role involves working with distributed systems, cloud environments, and high-throughput data pipelines on a daily basis. The engineer will collaborate with cross-functional teams to ensure reliable, scalable operations and contribute to architectural decisions.
Key facts
What you'll do
Maintain and monitor production systems that handle weather data ingestion and distribution at scale.
Design and implement infrastructure automation using configuration management and deployment tooling.
Troubleshoot performance issues across multi-tier applications serving real-time weather forecasts and alerts.
Collaborate with software engineers to containerize services and manage orchestration platforms.
Develop runbooks and operational procedures for incident response and system recovery.
Conduct capacity planning and performance tuning for databases, message queues, and caching layers.
Participate in on-call rotations and provide Tier 2 and Tier 3 support for production services.
Evaluate new technologies and present recommendations for improving system reliability and efficiency.
Write integration tests and validate deployments in staging environments before production rollout.
Document system architecture decisions and maintain up-to-date diagrams of the infrastructure topology.
Partner with data engineering teams to optimize pipelines for weather model output processing.
Configure and manage monitoring dashboards to track system health and service-level objectives.
Review and approve infrastructure changes through a structured change management and peer review process.
Contribute to disaster recovery planning and execute regular failover drills to validate system resilience.
Assist with the migration of legacy on-premise systems to cloud-based infrastructure and modern architectures.
Coordinate with security teams to implement access controls and audit logging across all environments.
Requirements
Bachelor's degree in Computer Science, Engineering, or a related technical discipline from an accredited institution.
Three or more years of hands-on experience administering Linux-based systems in production environments.
Demonstrated proficiency with Python or Go for writing scripts and automation utilities.
Practical experience with cloud platforms such as Amazon Web Services or Google Cloud Platform.
Solid familiarity with container technologies including Docker and Kubernetes for workload management.
Strong understanding of networking fundamentals, TCP/IP protocols, DNS resolution, and load balancing concepts.
Working knowledge of infrastructure-as-code tools such as Terraform or AWS CloudFormation for provisioning.
Ability to thrive in a fast-paced environment and effectively manage multiple concurrent priorities.
Experience with version control systems and collaborative code review practices in a team setting.
Excellent written and verbal communication skills for documenting systems and coordinating with stakeholders.
Nice to have
Prior experience with time-series databases such as InfluxDB or TimescaleDB for storing and querying metrics.
Background in meteorological, geospatial, or environmental data processing and analysis workflows.
Working knowledge of CI/CD pipelines using Jenkins, GitHub Actions, or similar automation tools.
Familiarity with observability platforms including Prometheus, Grafana, or Datadog for system monitoring and alerting.
Exposure to serverless computing patterns and event-driven architecture on cloud platforms.
Understanding of security best practices including IAM policies, network segmentation, and encryption at rest.
Skills & tools
Linux system administration and shell scripting for daily operations and troubleshooting.
Python or Go programming for automation, tooling development, and utility scripts.
AWS or Google Cloud Platform services and APIs for cloud infrastructure.
Docker containerization and Kubernetes orchestration for managing deployed workloads at scale.
Terraform or CloudFormation for infrastructure-as-code provisioning and configuration management.
Git version control and collaborative development workflows across distributed engineering teams.
SQL and NoSQL databases including PostgreSQL, MySQL, or Redis for data persistence.
RESTful APIs and HTTP protocols for service-to-service communication and integration.
Practical notes
This position is based in Golden, Colorado, and requires regular on-site presence at the company office.
The team follows a standard business schedule with some flexibility for asynchronous collaboration across time zones.
Candidates should expect to participate in on-call rotations on a periodic basis to support production systems.
The hiring process includes technical interviews, a systems design exercise, and a cultural fit conversation with the team.
The company offers competitive benefits including health insurance, retirement contributions, and professional development opportunities.
Applicants should be prepared to discuss past infrastructure projects and demonstrate problem-solving abilities during the interview.