Senior Cloud Engineer
Job description
About the role
Gridware is a technology company based in San Francisco that focuses on protecting and enhancing the electrical grid. The organization created a new discipline named active grid response (AGR). This discipline centers on monitoring electrical, physical, and environmental grid aspects that impact reliability and safety. Gridware uses a proprietary platform to deploy high-precision sensors for early issue detection. This method enables proactive maintenance and fault mitigation to reduce outages. The company's solutions aim to improve safety and ensure efficient grid operations. Financial backing comes from climate-tech and Silicon Valley investors. Further information is available on the official website.
The company is currently expanding its infrastructure for critical monitoring devices. These devices detect real-world fault events that can lead to wildfires. The role involves building and operating the platform that ingests millions of events daily. This system manages raw data streams and powers customer dashboards and alerting mechanisms. It also supports data science initiatives that convert raw signals into grid intelligence. The position requires ownership of AWS infrastructure, Kubernetes (EKS), CI/CD pipelines, and observability systems. The candidate will work closely with the Cloud Security team. Collaboration with backend, firmware, and data teams is essential to maintain deployment velocity. This position is for an early DevOps team member. The work will directly influence how the company builds and runs production systems.
The primary responsibility is designing and running the ingestion pipeline for high-volume sensor data. This platform moves field signals into operational dashboards and alerts. It maintains support for scientific analysis teams. The candidate will partner with firmware, data, and security groups. The goal is to keep deployments fast, compliant, and reliable for grid infrastructure.
Specific duties include designing intake mechanisms for reliable high-volume data streaming. The engineer will build Kubernetes patterns on AWS to host monitoring services. Ensuring stability for critical grid monitoring services is required. The role involves reviewing code and configurations for safety and auditability. Compliance with standards must be maintained rigorously. The engineer will ship observability tools. These tools provide backend and firmware teams with rapid insight into device behavior. Partnership with security groups is necessary to harden pipelines and access rules. Construction of CI/CD routes allows small backend teams to release frequently. Reliability must remain high during these deployments. Refactoring infrastructure definitions reduces manual steps and increases consistency. Runtime behavior is guarded by tuning alert thresholds and dashboards. The engineer responds to emerging fault patterns in the grid.
The requirements specify three to five years of managing cloud infrastructure. Production experience with Kubernetes is mandatory. Understanding of networking and storage patterns is essential. Knowledge of Linux systems that run critical services is required. The candidate must write infrastructure as code. These tools must integrate with automated testing frameworks. Clear communication skills are necessary for collaboration. Engineers build firmware, data models, and security controls.
Nice to have qualifications include open source contributions. Experience with distributed systems projects is valued. Specific skills include AWS, Kubernetes, EKS, CI/CD, Terraform, Python, and Grakana.
Practical application details require confirmation Compensation details specify a salary of $180,000 per year. The engagement type is Full-Time. The location for this role is San Francisco, CA.
What you'll do
- Architect and maintain the ingestion pipeline that processes millions of events per second from grid monitoring sensors.
- Build and operate Kubernetes (EKS) patterns on AWS to host critical monitoring services for the Gridware platform.
- Ensure stability and high availability of the platform that drives customer dashboards and alerting mechanisms for grid operators.
- Design intake mechanisms for reliable high-volume data streaming to support real-time grid intelligence.
- Partner with firmware, data, and security teams to keep deployments fast, compliant, and reliable for critical grid infrastructure.
- Review code and configurations with a focus on safety, auditability, and compliance with industry standards.
- Ship observability tools that provide backend and firmware teams with rapid insight into field device behavior and system health.
- Harden pipelines and access rules in collaboration with the Cloud Security team to protect data and control pathways.
- Construct and optimize CI/CD routes that allow small backend teams to release frequently without compromising reliability.
- Refactor infrastructure definitions to reduce manual steps, increase consistency, and streamline operations.
- Tune alert thresholds and dashboards to guard runtime behavior and respond to emerging fault patterns in the grid.
- Support scientific analysis teams by maintaining the data streams and storage patterns that power long-term grid intelligence.
- Implement robust monitoring and logging solutions to ensure operational stability and rapid troubleshooting.
- Participate in on-call rotations to address incidents and ensure continuous availability of critical grid monitoring services.
Requirements
- Bring three to five years of experience managing cloud infrastructure in a production environment.
- Possess mandatory production experience with Kubernetes and a deep understanding of its operational nuances.
- Demonstrate a strong grasp of networking and storage patterns as they apply to distributed systems.
- Show knowledge of Linux systems that run critical services and how to maintain them securely.
- Write infrastructure as code using modern tools that integrate with automated testing frameworks.
- Collaborate effectively with cross-functional teams including firmware, data, and security engineers.
- Communicate clearly and constructively during design discussions, incident reviews, and planning sessions.
- Build solutions that meet the stringent safety and compliance needs of grid infrastructure monitoring.
- Contribute to a culture of reliability, auditability, and performance in all deployed services.
Nice to have
- Contribute open source contributions that demonstrate engineering rigor and collaboration.
- Engage with distributed systems projects that solve challenging problems at scale.
- Apply specific skills with AWS, Kubernetes, EKS, CI/CD, Terraform, Python, and Grafana.
Practical notes
- Compensation for this role is $180,000 per year.
- Engagement is Full-Time.
- Location for this role is San Francisco, CA.