AI Infrastructure Engineer
FinUK3w ago
Job description
About the role
Fin is on the lookout for experienced engineers to enhance the systems that underpin our proprietary AI models. This role involves joining a dedicated team that focuses on the entire spectrum of AI infrastructure, from optimizing GPU performance to deploying advanced models such as Fin Apex. You will play a crucial part in shaping the future of our AI capabilities.
Key facts
What you'll do
- Architect and scale distributed training pipelines specifically designed for large language models (LLMs) and transformer architectures, encompassing everything from data preprocessing to final model evaluation.
- Develop and sustain inference services that emphasize low latency and high reliability, incorporating features such as autoscaling and effective traffic management.
- Enhance GPU performance by diagnosing performance bottlenecks and fine-tuning computational kernels for optimal efficiency.
- Collaborate closely with machine learning scientists to transition experimental training and inference methodologies into robust production systems.
- Provide mentorship to team members, fostering a culture of operational excellence and ensuring high standards of system reliability.
- Engage in code reviews and contribute to best practices in software development within the team.
- Participate in the design and implementation of monitoring and alerting systems to ensure the health and performance of AI infrastructure.
- Stay updated with the latest advancements in AI and machine learning technologies, integrating relevant innovations into existing systems.
Requirements
- A minimum of 5 years of experience in software engineering, demonstrating a strong technical foundation.
- A degree in Computer Science, Computer Engineering, or a related field, or equivalent hands-on experience in the industry.
- Proficiency in at least one programming language, such as Python, Ruby, Java, or Go, with a solid understanding of software development principles.
- Practical experience in at least one of the following areas: training models for transformers or LLMs, executing model inference at scale, or low-level GPU programming using CUDA or Triton.
- Familiarity with production environments, particularly those that operate at scale, is essential.
- Strong problem-solving skills and the ability to work collaboratively in a team-oriented environment.
Nice to have
- Experience working at AI-focused companies that handle their own model training or inference processes.
- Knowledge of Kubernetes for workload management and orchestration.
- Familiarity with cloud services, particularly AWS or other leading cloud platforms.
- Experience using Python for infrastructure or machine learning tasks in a production setting.
- Involvement in open-source projects, personal technical endeavors, or contributions to published materials in the field.
Skills & tools
- Expertise in transformers and large language models (LLMs)
- Proficient in CUDA and Triton for GPU programming
- Programming skills in Python, Ruby, Java, or Go
- Experience with Kubernetes for container orchestration
- Familiarity with AWS or similar cloud services
- Knowledge of distributed training and inference systems
Practical notes
- The compensation package includes a competitive salary along with equity options.
- Benefits encompass a pension plan with a matching contribution of up to 4%, comprehensive health and dental coverage, life insurance, and a cycle-to-work initiative.
- Flexible paid time off is available to support work-life balance.
- Parental leave benefits include paid maternity leave and six weeks of paid paternity leave.
- Daily lunches, snacks, and a fully stocked kitchen are provided to employees.
- Each employee is equipped with a MacBook, with Windows options available for specific roles.
- Access to Claude Code and other AI tools is provided to encourage experimentation and innovation.
Join us at Fin and be part of a dynamic team that is at the forefront of AI infrastructure development. Your expertise will help drive the evolution of our AI capabilities, making a significant impact on our products and services. If you are passionate about AI and eager to work in a collaborative environment, we would love to hear from you.