ML Infrastructure Engineer
Job description
About the role
Senior ML Infrastructure Engineer
Gridware maintains its headquarters in San Francisco, California. The company operates a full-time position for a Senior ML Infrastructure Engineer within its Automation organization. This role reports directly to the automation leadership and works alongside machine learning, operations, and analytics teams. The primary mission involves scaling the time-saving benefits delivered to customers through enhanced infrastructure for model deployment and monitoring. You will architect the foundational systems that translate complex grid sensor data into actionable intelligence for utility operators. This position demands ownership of the end-to-end lifecycle responsible for moving models from experimental notebooks to robust production services. Your daily decisions will directly influence the stability and responsiveness of the platform that powers active grid response. Success in this role directly accelerates the value Gridware provides to utility partners by ensuring infrastructure keeps pace with innovation.
Key facts
What you'll do
- Execute the work described for as detailed
- Architect and maintain the containerized infrastructure that hosts machine learning models for grid analytics.
- Construct intake pipelines that normalize diverse sensor readings and signal data before models consume the information.
- Establish and enforce review checkpoints that quantitatively measure model performance against historical grid events and operational baselines.
- Coordinate release procedures to ensure new deployments integrate seamlessly with existing monitoring, alerting, and logging systems.
- Implement monitoring dashboards that provide deep visibility into inference latency, data quality metrics, and system resource utilization.
- Forge partner alignments that synchronize deployment schedules with utility operational calendars and maintenance windows.
- Refine data retention policies to balance strict regulatory compliance with practical storage limitations and cost efficiency.
- Optimize resource allocation strategies to improve the efficiency and throughput of training jobs on shared compute clusters.
- Build robust containerized environments where automated tests continuously validate the integrity of grid data workloads.
- Ensure model hosting platforms are reliable, observable, and provide the necessary telemetry for the platform team.
- Drive the standardization of infrastructure definitions to enable consistent and repeatable deployments across environments.
- Collaborate closely with analytics and operations groups to resolve complex infrastructure bottlenecks impacting model performance.
- Safeguard platform integrity by implementing security and compliance practices tailored to the critical nature of grid data.
- Proactively identify infrastructure risks and lead remediation efforts to minimize downtime and service disruptions.
Requirements
- Must possess hands-on experience with Kubernetes and container orchestration platforms to manage production workloads.
- Must have a minimum of three years managing model deployment in production settings within complex software environments.
- Must demonstrate proficiency with Python codebases used for infrastructure automation and configuration management.
- Must have a deep understanding of monitoring concepts specific to distributed systems and microservices architectures.
- Must have direct experience processing time series data originating from sensors, meters, or IoT devices in industrial settings.
- Must be familiar with grid operations terminology to facilitate clear communication with utility engineering teams.
- Must be able to work full-time in San Francisco, CA, as specified by the engagement details.
- Must submit application details through the official apply page to be considered for this specialized role.
Nice to have
No specific nice-to-have qualifications are currently specified for this position.
About the company
Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid. We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the electrical, physical, and environmental aspects of the grid that affect reliability and safety.