AI Engineer, AI & Applications
Job description
This page establishes the role of AI Engineer within the AI & Applications function at Firmus Technologies. The position is responsible for constructing the foundational production workflows of Firmus AI Factory. Success in this role ensures efficient, scalable distributed training capabilities for current and future model development.
About the role
The AI Engineer establishes the operational backbone of Firmus AI Factory. This position designs and delivers production-grade training systems that scale efficiently. The work involves creating pre-built training recipes, evaluation benchmarks, and model guidance that become the standard reference for all hyperscale customer workflows. The role serves as a critical bridge between platform development and customer implementation.
Key facts
Location options include Singapore or Australia, with specific locations such as Launceston, TAS or Sydney, NSW. The engagement is full-time. The base compensation is set at $180,000 SGD for the year 2024.
What you'll do
You will design intake procedures that validate data quality and training requirements before any training work commences. The role requires orchestrating build processes for major frameworks including TorchTitan and Megatron-LM. You will manage parallelism strategies spanning FSDP and tensor parallel implementations.
A core responsibility is guiding review cycles using rigorous benchmarks. These benchmarks track accuracy, latency, token cost, and other critical performance indicators. You will champion ship rituals that harden model checkpoints and streamline deployment processes on Kubernetes and Slurm clusters.
The position involves forging deep partner integrations. This alignment connects job scheduling systems with model optimization efforts for both inferencing and application workloads. You will conduct fine-tuning experiments, specifically exploring LoRA and QLoRA methodologies. The goal is to generate measurable gains for operations domain data.
You will create and publish efficiency playbooks. These documents enable customers to optimize their workloads and reduce configuration errors. The role also requires championing evaluation harnesses. These systems power leaderboard tracking and facilitate standardized model comparison across teams.
Driving benchmarking suites is a primary duty. These suites must measure throughput, energy efficiency, and model flops utilization (MFU). You will facilitate cross-team synchronization. This ensures clear communication patterns between AI engineers and software engineering staff regarding optimization trade-offs.
Requirements
The candidate must bring 5-7 years of hands-on experience in distributed machine learning. Demonstrated expertise must exist in PyTorch and JAX environments across many GPUs. A deep understanding of GPU optimization is essential. This includes utilization metrics, memory allocation patterns, and NCCL collective communication behavior.
Practical experience debugging convergence issues and profiling bottlenecks is required. The candidate must have optimized throughput in large-scale training runs exceeding standard configurations. Designing controlled experiments and communicating results with rigorous methodology is mandatory.
Familiarity with production frameworks is required. This includes TorchTitan, Megatron-LM, and similar training stack configurations. The candidate must clearly explain model parallelism trade-offs. These trade-offs exist between FSDP, tensor parallelism, pipeline strategies, and other architectural options.
Key competencies
The role demands Distributed Systems Mastery. The engineer must explain NCCL operations, collective communications, and the root causes of scaling inefficiency. Benchmarking Rigor is another core competency. The engineer must validate assumptions, explain variance, and communicate uncertainty in results.
Production Thinking is essential. The engineer must understand checkpointing, recovery procedures, resource constraints, and cost optimization strategies. Mentorship capability is required to guide other engineers on training best practices and debugging distributed issues. Excellence in Documentation ensures the creation of clear, actionable playbooks that customers can reliably follow.
Success metrics
The effectiveness of the role will be measured by specific outcomes. Benchmark credibility and decision impact must increase. Benchmarks should be trusted and actively used to drive model, hardware, and product decisions. Training efficiency leadership is expected, shown by sustained improvement in benchmarked training efficiency.
The time-to-validate new models must shorten, allowing for quick and consistent end-to-end evaluation. Template effectiveness should improve, reducing misconfigurations and repeated setup failures. This leads to fewer training config escalations. The final metric is Competitive differentiation, where model arena outputs directly influence customer adoption and internal roadmap priorities.
Location and reporting
The position is based in Singapore or Australia, with work locations in Launceston, TAS or Sydney, NSW. The role reports directly to the Head of AI & Applications.
Employment basis
This is a full-time position.
Diversity
Firmus is committed to building a diverse and inclusive workplace. The company encourages applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engine