Member of Technical Staff
Job description
Member of Technical Staff at Emerald Ai.
About the role
You will design and implement real-time telemetry pipelines that capture high-frequency data from power meters, PDUs, UPS systems, and cooling infrastructure across distributed datacenter environments. This role requires you to integrate directly with on-site industrial power systems and the electrical grid, translating raw operational metrics into actionable control signals. You will own the architecture for fault-tolerant integration layers that connect IT systems like cloud APIs and databases with OT systems such as SCADA and DCIM. Your work will ensure safe and reliable behavior of control systems, especially during degraded, disconnected, or split-brain network states where safety and correctness are paramount. As an early engineering hire, you will shape system architecture, define data model mappings, and own projects from initial design through production deployment. You will collaborate closely with cross-functional teams to translate real-world energy and facility constraints into scalable software solutions. Your contributions will directly determine how dynamically workloads can be shifted to align with grid conditions and renewable energy availability. Ultimately, you will play a critical role in enabling AI data centers to scale without overwhelming the electrical grid.
Key facts
What you'll do
- Build software modules that interact with datacenters, on-site industrial power systems, and the electrical grid to enable dynamic load management.
- Architect and implement data ingest pipelines that collect, normalize, and persist high-frequency telemetry from power meters, PDUs, UPS systems, cooling infrastructure, and compute hardware.
- Design fault-tolerant integration layers between IT systems such as cloud APIs, databases, and orchestration platforms and OT systems, handling protocol translation and data model definition and mapping.
- Ensure high safety, reliability, and correctness when interacting with real-world energy assets and critical facilities, including behavior in degraded, disconnected, or split-brain network states.
- Collaborate across teams to deliver end-to-end solutions from edge device to control plane, ensuring consistent operation under varying network and infrastructure conditions.
- Implement control and optimization logic that dynamically adjusts workloads based on grid signals, infrastructure telemetry, and operational policies.
- Partner with data scientists and product teams to incorporate energy market signals, demand response logic, and optimization strategies such as load shifting and peak shaving into system behavior.
- Define and maintain operational dashboards, alerting, and observability pipelines that provide clear insight into power, compute, and cooling interactions.
- Write robust integration tests, simulate failure modes, and validate behavior under constrained or emulated grid and facility conditions.
- Contribute to system design reviews, code reviews, and architectural documentation to ensure long-term maintainability and scalability.
- Work closely with field engineers and customers to understand real-world constraints and translate them into software requirements and edge deployment strategies.
- Support production incidents, perform root cause analysis, and implement improvements that increase resilience of telemetry and control paths.
- Mentor junior engineers by providing guidance on best practices for building reliable, high-performance telemetry and control systems.
- Help define the roadmap for energy-aware control features, contributing to decisions on protocol support, data model evolution, and integration patterns.
Requirements
- 7+ years of software engineering experience, with strong proficiency in Python and/or Rust, Go or C/C++.
- Hands-on experience building telemetry ingest pipelines or distributed systems that handle high-volume, time-series data from sensors and control devices.
- Familiarity with at least one async messaging system such as Kafka, RabbitMQ, or equivalent platforms for reliable data transport.
- Exposure to IT or OT systems including SCADA, EMS, BMS, or DCIM, and understanding of how data flows between these environments.
- Ability to work safely in environments where software errors can impact physical infrastructure, recognizing the critical nature of reliability and correctness.
- Experience designing systems that must operate correctly during network partitions, device disconnections, or other degraded states.
- Strong understanding of networking concepts, protocols, and serialization formats commonly used in industrial and datacenter environments.
- Commitment to following defined processes for safety, change management, and operational procedures in critical infrastructure settings.
Nice to have
- Direct experience with industrial telemetry protocols such as Modbus TCP/RTU, DNP3, OPC-UA, BACnet, or SNMP.
- Familiarity with time-series databases like InfluxDB, TimescaleDB, Prometheus, or industrial historian platforms such as AVEVA PI or OSIsoft, and Ignition.
- Experience with Kubernetes, containerized deployments, or HPC job scheduling in datacenter environments.
- Background in power systems, energy markets, microgrids, or datacenter infrastructure including power distribution, cooling, UPS, and PDUs.
- Understanding of control system design patterns such as PID loops, state machines, setpoint control, or demand response logic.
- Experience with edge compute runtimes or constrained environments where low-latency, reliable local execution is essential when cloud connectivity is intermittent.
- Familiarity with optimization algorithms or energy management strategies such as load shifting, peak shaving, and curtailment.
Practical notes
- Bay Area location with 1 WFH day per week.
- Full-time engagement with competitive compensation and equity.
- Employment is contingent on eligibility to work in the United States.
- The company is an equal opportunity employer and makes reasonable accommodations for applicants with disabilities.