Site Reliability Engineer, Storage - Enterprise Technology
Job description
About the role
In this role, you will play a vital role in maintaining and enhancing the reliability of our storage systems. You will work closely with cross-functional teams to ensure high availability and performance of our infrastructure, contributing to the overall efficiency of our trading operations. Your expertise will be instrumental in driving initiatives that improve data durability and access speeds across our enterprise environment. You will own the design and execution of strategies that mitigate risks associated with storage failure scenarios. This position requires a proactive approach to identifying potential infrastructure weaknesses before they impact critical systems. You will be a key contributor to the continuous improvement of our operational processes and standards. Your work will directly support the firm's trading activities by providing a stable and performant underlying storage platform. Through collaboration with engineering partners, you will help shape the future of our technological infrastructure.
Key facts
What you'll do
- Architect and implement robust storage solutions that align with the high demands of financial trading environments.
- Conduct in-depth performance analysis to identify bottlenecks and drive optimization initiatives across storage infrastructures.
- Engineer sophisticated automation frameworks to eliminate manual overhead and reduce the potential for operational error.
- Establish comprehensive monitoring dashboards and alerting mechanisms to ensure immediate visibility into system anomalies.
- Lead post-incident reviews to dissect complex system failures and formulate actionable preventative measures.
- Partner with development teams to integrate storage best practices into the software development lifecycle from the outset.
- Evaluate and deploy emerging storage technologies to maintain a competitive edge in infrastructure capabilities.
- Coordinate with network operations to ensure seamless connectivity and data flow between storage arrays and application servers.
- Document complex system architectures and operational procedures to maintain institutional knowledge and facilitate onboarding.
- Champion the adoption of infrastructure as code methodologies to ensure consistency and repeatability in deployments.
- Manage backup and disaster recovery strategies to safeguard critical financial data and ensure business continuity.
- Interface with vendors and stakeholders to resolve complex technical issues related to storage hardware and software.
- Analyze capacity trends to guide future infrastructure investments and prevent resource constraints.
- Serve as a technical leader within the storage domain, mentoring junior engineers and elevating team capabilities.
Requirements
- Hold a Bachelor's degree in Computer Science, Engineering, or a related field from an accredited institution.
- Possess a minimum of 3 years of professional experience in site reliability engineering or a functionally similar role within a high-scale technical environment.
- Demonstrate proficiency in Linux/Unix systems administration and advanced scripting capabilities using languages such as Python or Bash.
- Exhibit a strong, comprehensive understanding of enterprise storage technologies, including Storage Area Networks (SAN), Network Attached Storage (NAS), and modern cloud storage solutions.
- Have hands-on experience with system monitoring tools and methodologies for performance tuning and capacity planning.
- Show familiarity with containerization platforms and orchestration systems such as Docker and Kubernetes as they apply to storage workloads.
- Maintain a proven track record of working effectively within a collaborative, cross-functional team structure common in fast-paced trading firms.
- Possess strong analytical and problem-solving skills to troubleshoot complex issues under pressure and with tight operational timelines.
- Understand the principles of data integrity, security, and compliance as they relate to enterprise storage management.
- Demonstrate the ability to manage multiple priorities and technical tasks simultaneously in a dynamic and demanding environment.
- Have experience with storage protocols, data replication techniques, and network storage configurations.
- Commit to adhering to established operational procedures and contributing to a culture of safety and reliability.
- Be prepared to participate in on-call rotations to address critical infrastructure issues as they arise during business hours.
- Hold authorization to legally work in the United States without sponsorship requirements for the position advertised.
Nice to have
- Acquire knowledge of financial trading systems and market infrastructure to better align storage strategies with business objectives.
- Utilize infrastructure as code tools such as Terraform or Ansible to automate the provisioning and management of storage resources.
- Apply familiarity with databases and data management practices to optimize storage interactions for transactional workloads.
Practical notes
This role requires candidates to be authorized to work in the United States. Applications will be accepted until the position is filled. Hudson River Trading is an equal opportunity employer and encourages all qualified candidates to apply. The company provides a comprehensive benefits package designed to support the well-being and professional growth of its employees.