Member of Technical Staff
Job description
Member of Technical Staff at Emerald Ai.
About the role
You will design and build novel systems that make AI infrastructure operate intelligently under power constraints, directly addressing the critical bottleneck of escalating compute demand on the electrical grid. This role centers on inventing and shipping software that dynamically adjusts data center behavior to prevent overloads while enabling massive AI scale. You will operate at the intersection of distributed systems, machine learning, cloud infrastructure, and energy-aware optimization to move research prototypes into robust production deployments. The position is deeply hands-on, requiring you to build production-quality infrastructure and large-scale experimental platforms that manage thousands of GPUs in real environments. You will work closely with product and customer teams to translate novel ideas into deployed capabilities that customers can rely on for their most demanding workloads. A core part of your mission will be to publish cutting-edge research that advances the state of the art in AI infrastructure and energy-efficient computing. You will also play a key role in shaping Emerald AI's long-term technical roadmap and identifying new research directions with clear commercial impact.
Key facts
What you'll do
- Design and build novel systems for power-aware AI infrastructure, distributed computing, and large-scale cloud platforms.
- Develop production-quality software, research prototypes, and experimental infrastructure that can be deployed in real-world AI data centers.
- Apply machine learning, optimization, systems, or control techniques to challenging problems in AI infrastructure and cloud operations.
- Design, implement, and evaluate algorithms using large-scale experimental platforms and production deployments.
- Partner with product and customer facing teams to transition research innovations into customer-facing products.
- Collaborate with partners across industry and academia on cutting-edge research initiatives.
- Publish high-impact research in leading systems and AI conferences when appropriate.
- Help shape Emerald AI's long-term technical roadmap and identify new research directions with commercial impact.
- Take ownership of the entire lifecycle for critical infrastructure components from initial design through deployment and iteration.
- Diagnose and resolve complex performance and stability issues in production systems that span multiple data center sites.
- Implement observability and debugging tooling to ensure reliability and transparency for internal platform users.
- Optimize software stacks to reduce latency, increase throughput, and minimize energy consumption across the compute fleet.
- Define and drive experiments that validate system behavior under realistic AI training and inference workloads.
- Work closely with hardware teams and partners to maximize utilization of accelerators and power delivery infrastructure.
Requirements
- Hold a Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field.
- Machine learning systems
- AI infrastructure
- Distributed systems
- Cloud computing
- Systems for AI or HPC
- Performance optimization
- Demonstrate excellent software engineering skills with extensive experience developing large software systems in languages such as C++, Python, Go, or Rust.
- Have a proven track record of building research prototypes and scaling them to large-scale production codebases.
- Maintain a strong publication record or a demonstrated history of delivering impactful technical innovations.
- Show the ability to independently drive research from initial concept through implementation and rigorous evaluation.
- Exhibit strong analytical and problem-solving skills when working on ambiguous, complex, and high-stakes technical challenges.
- Communicate effectively in written and verbal form with both technical and non-technical stakeholders.
Nice to have
- Experience deploying systems in production cloud or distributed environments.
- Experience working with large codebases, production software, or open-source infrastructure.
- Experience with Kubernetes, Slurm, distributed training/inference frameworks, or large-scale AI infrastructure.
- Experience with GPU systems, accelerators, or performance analysis tools.
- Experience in optimization, control systems, resource scheduling, or systems performance.
- Experience with power management, energy-efficient computing, sustainability, or data center infrastructure.
- Experience taking research innovations from prototype to production.
Practical notes
- Location is in the Bay Area.
- Engagement is full_time.
- The role offers flexibility with 2 WFH days per week.
- Compensation includes competitive pay plus equity; comprehensive benefits cover medical, dental, vision, and 401(k) matching.
- The position is not eligible for remote work outside the Bay Area.
- The candidate must be available to work onsite in the Bay Area with the option for limited remote days.
- Employment is at-will and contingent on standard background and reference checks.
- This role requires the ability to collaborate effectively on-site with cross-functional teams during core working hours.