Technical Lead, Machine Learning
Job description
About the role
Join the foundational engineering team for a new AI venture focused on building next-generation AI-native productivity tools. This role is for an individual who can translate advanced AI research into stable, production-ready systems. You will own the architecture and delivery of core machine learning capabilities that define the product experience. The position requires deep collaboration with product and research to align technical feasibility with ambitious feature goals. You will be responsible for establishing best practices across the ML lifecycle from experimentation to deployment. This role is critical for setting the technical direction of the AI stack. Your work will directly influence the reliability and scalability of the platform from day one. You will mentor engineers and standardize workflows to ensure consistent execution across the team.
What you'll do
Architect and deliver the end-to-end machine learning pipeline for the AI-native productivity suite. Design and implement scalable training workflows utilizing PyTorch and JAX for next-generation model development. Lead the integration of inference optimization strategies to ensure low-latency performance in production environments. Define and manage the ML infrastructure stack, including data processing, feature stores, and experiment tracking. Collaborate closely with research partners to iterate on model architectures and align on performance benchmarks. Own the evaluation framework for models, establishing rigorous metrics for quality and operational efficiency. Drive the adoption of distributed training techniques to accelerate model development cycles. Oversee GPU system provisioning and resource allocation to support efficient model training and testing. Implement robust fine-tuning procedures that adapt base models to specific product requirements and user needs. Establish monitoring and observability tools to track model behavior and system health in live deployments. Lead technical design discussions and make informed trade-offs between accuracy, speed, and maintainability. Champion code quality and documentation to ensure knowledge sharing and long-term maintainability. Guide the selection and integration of tools for model versioning, deployment automation, and testing. Mentor and uplevel junior engineers through code reviews, technical guidance, and hands-on pair programming. Translate high-level product objectives into detailed technical specifications for the ML team.
Requirements
Demonstrated expertise with Python, showcasing advanced proficiency in building scalable and maintainable codebases. Extensive experience with PyTorch, including model definition, training loops, and optimization techniques. Strong background in JAX, with a clear understanding of its functional programming paradigm and performance benefits. Hands-on experience with LLM fine-tuning methods, including supervised fine-tuning and reinforcement learning from human feedback. Deep knowledge of inference optimization, covering model quantization, pruning, and efficient serving strategies. Solid understanding of ML infrastructure, including data pipelines, feature engineering, and experiment management. Proven track record with distributed training across multiple GPUs and nodes to handle large-scale model workloads. Comprehensive experience with GPU systems, including driver management, memory optimization, and hardware utilization. Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field. Fluency in English is required for effective collaboration with distributed teams and stakeholders. Minimum of eight years of professional experience in software engineering or machine learning roles. Demonstrated ability to operate in a fast-paced environment with minimal supervision and high ownership. Strong problem-solving skills and a track record of debugging complex issues in production systems. Commitment to writing clean, tested, and maintainable code that adheres to industry best practices.
Nice to have
Experience with A1's first product, which reimagines email with autonomous AI agents. Contributions to open-source ML projects or publications in relevant conferences. Deep understanding of financial systems and trading concepts. Experience with multi-cloud environments and container orchestration platforms. Knowledge of compliance and regulatory considerations in financial technology. Familiarity with broker-dealer operations and risk management systems. Exposure to emerging markets and regional technology adoption trends.
Practical notes
This role is part of A1, an AI venture incubated by BJAK, with an initial $100M investment. The team is small, high-calibre, and focused on complex AI infrastructure challenges. The company operates remotely with a focus on long-term career development.