Storage and Datacenter Team Manager
Job description
About the role
You will lead and architect the storage and datacenter infrastructure that powers Voleon's AI and machine learning driven investment strategies. In this role, you own the design, reliability, and scalability of the systems that store and protect our critical financial data. You will provide hands-on technical leadership by implementing complex storage architectures and solving difficult infrastructure challenges. You will mentor and grow a team of engineers, fostering a culture of excellence and continuous improvement. You will drive automation to streamline operations and reduce manual toil across the storage environment. This position requires deep expertise in managing data lifecycle processes, including archiving and backup at scale. You will work closely with world-class researchers and technical professionals in a collaborative and high-performance environment. Success in this role means ensuring that our infrastructure is robust, efficient, and aligned with ambitious business objectives.
Key facts
What you'll do
Lead a small team of storage, database, and systems administrators with duties including mentorship, performance management, and career development.
Coordinate datacenter operations across production and research facilities, including scheduling site work and managing vendor and contractor visits.
Align team priorities with organizational goals and ensure timely delivery of projects to support the business.
Participate in hiring efforts to grow and evolve the storage engineering team with high-caliber technical professionals.
Coordinate on-call schedules and ensure effective incident response processes are in place for storage and datacenter events.
Architect, implement, and maintain highly available and performant storage systems that meet the demands of a multibillion-dollar asset manager.
Define and drive automation strategies for storage deployment, monitoring, and operations to improve efficiency and scalability.
Oversee storage lifecycle management including capacity planning, performance tuning, and data protection strategies such as archiving and backups for large-scale datasets.
Provide architectural guidance and hands-on support for Ceph storage systems operating at petabyte scale.
Oversee physical datacenter infrastructure including rack layout planning, power distribution, cooling systems, and capacity forecasting for space, power, and cooling.
Manage equipment installation, upgrades, and decommissioning in a secure and efficient manner.
Collaborate closely with networking, virtualization, research, and application teams to support diverse compute and storage needs across the organization.
Participate in and improve CI/CD and configuration management processes with tools like Ansible and Git to ensure consistency and reliability.
Support database operations through database tuning, storage optimization, and collaboration with developers to maintain high performance.
Develop and maintain runbooks for remote-hands work and coordinate with contractors and facility personnel to perform necessary onsite operations.
Serve as an escalation point for advanced troubleshooting of distributed filesystems, databases, and high-performance storage infrastructure.
Perform hands-on administration of Linux servers, network-attached storage, virtualization platforms, and cluster frameworks as needed.
Install, cable, and troubleshoot physical server, storage, and network hardware in demanding rack environments.
Diagnose and resolve hardware-level issues impacting production systems to maintain uptime and reliability.
Support and enhance observability using tools such as Prometheus, Grafana, and other monitoring platforms to ensure system health.
Requirements
Bring 5 or more years of Linux Systems Administration experience with a significant recent focus on storage systems and technologies.
Demonstrate 2 or more years of team leadership, technical project management, or mentoring experience guiding technical professionals.
Show deep knowledge of distributed storage systems such as Ceph and understand storage technologies including RAID, SAN, and NAS.
Have hands-on experience streamlining data lifecycle processes, including archiving, backup, and retention of PB-scale data in demanding environments.
Possess direct experience with co-located data center infrastructure and the operational realities of facility management.
Be prepared to travel to remote datacenter sites when needed to address issues and perform strategic work.
Show strong scripting and development experience in Bash and or Python to automate complex tasks and solve operational problems.
Have practical experience with configuration management tools like Ansible to implement infrastructure as code.
Demonstrate familiarity with monitoring and alerting systems such as Nagios CheckMK, Prometheus, and Grafana for operational visibility.
Have a solid understanding of virtualization technologies including KVM and ESXi, as well as containerization platforms like Docker and Podman.
Show knowledge of LDAP IPA and Active Directory for centralized identity management in enterprise environments.
Nice to have
Only pursue preferred qualifications if they align with your background and the demands of the role.
Practical notes
This is a full-time position based in Berkeley, California. Occasional participation in on-call rotation is expected as part of the operational responsibilities.