Member of Technical Staff - Compute Cluster
Job description
Member of Technical Staff - Compute Cluster at Causal.
About the role
Causal is at the forefront of developing a Large Physics Foundation Model (LPM) aimed at predicting and influencing various physical systems, with an initial focus on weather patterns. In this position, you will be responsible for constructing and maintaining the high-performance computing infrastructure that underpins our AI research and development initiatives. Your primary goal will be to ensure the efficient and reliable operation of our GPU fleet, which is critical for tasks such as training, evaluation, and serving AI models.
Key facts
What you'll do
- Architect, implement, and oversee extensive distributed GPU clusters, managing all aspects from provisioning to imaging, updates, and capacity planning.
- Enhance existing scheduling and orchestration frameworks, such as Kubernetes or Slurm, to optimize resource allocation, enable preemption, enforce quotas, and facilitate multi-user access for AI workloads.
- Develop software tools that streamline cluster management, creating a unified self-service platform for researchers and engineers to utilize.
- Oversee cluster storage solutions and data pathways for checkpoints and logs, ensuring that retention policies and data lineage are well-defined and adhered to.
- Continuously monitor and enhance system reliability and error recovery processes, creating observability tools that allow for proactive issue detection.
- Work closely with researchers to support large-scale computational tasks, providing insights and recommendations on performance optimization and resource allocation.
- Collaborate with cross-functional teams to identify and resolve infrastructure challenges, ensuring seamless integration with ongoing projects.
- Document processes, configurations, and best practices to facilitate knowledge sharing and onboarding of new team members.
Requirements
- Proven experience in managing large GPU clusters and familiarity with container orchestration systems like Kubernetes, Slurm, or Docker.
- A solid foundation in systems administration, with expertise in Linux, networking, storage solutions, and infrastructure-as-code methodologies.
- Proficiency with cloud service providers such as GCP, AWS, or Azure, particularly in relation to their machine learning and AI offerings.
- Understanding of industry best practices for monitoring, logging, observability, and version control within machine learning environments.
- Knowledge of CUDA and NCCL, along with techniques for profiling and optimizing performance in distributed computing scenarios.
- Demonstrated ability to take ownership of projects, managing them from the initial requirements phase through to independent execution and delivery.
Nice to have
- Experience with machine learning frameworks and libraries, such as TensorFlow or PyTorch.
- Familiarity with data engineering practices and tools for handling large datasets.
- Previous involvement in research projects, particularly in the AI or physics domains.
Skills & tools
- Kubernetes
- Slurm
- Docker
- Linux
- GCP
- AWS
- Azure
- CUDA
- NCCL
Practical notes
- This position is based in San Francisco, a vibrant tech hub that offers a wealth of opportunities for professional growth and networking.
- As a full-time role, you will be expected to engage with the team regularly and contribute to the collaborative environment that Causal fosters.
- While compensation details are not specified, candidates can expect a competitive salary commensurate with experience and industry standards.
- Causal is committed to supporting its employees through various benefits and opportunities for professional development.
- Visa sponsorship may be available for qualified candidates, making this an attractive opportunity for international applicants looking to work in the United States.
In summary, if you are passionate about high-performance computing and eager to contribute to groundbreaking AI research, we invite you to apply for the Member of Technical Staff position at Causal. Join us in shaping the future of predictive modeling and physical systems analysis.
About the company
Causality is an influence by which one event, process, state, or subject contributes to the production of another event, process, state, or object where the cause is at least partly responsible for the effect, and the effect is at least partly dependent on the cause. The cause of something may also be described as the reason behind the event or process.