Principal Solutions Architect, AI Data Infrastructure
Job description
About the role
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
In this context, the Principal Solution Architect, AI Data Infrastructure plays a critical role within the AI Cloud and AI Factory organizations. You will own the architecture, evaluation, and optimisation of the data systems that feed Firmus AI Cloud and the AI Factory. You will serve as Firmus's subject-matter expert on where data lives and moves across the AI fabric - spanning cache layers, memory fabrics, and high-performance storage systems that keep GPUs fully utilised. You will translate complex data-path decisions into measurable outcomes for performance, capacity, and total cost of ownership. In a business built on making every watt count, you will make every byte count with it. You will act as the trusted technical authority for both internal engineering teams and external customers, validating workload requirements and shaping our platform roadmap through deep, principles-based understanding.
The role is based in Singapore and operates as a full-time position.
What you'll do
Define the reference architecture for the AI data path end to end, including parallel and software-defined storage, NVMe and NVMe-oF, object and file systems, memory disaggregation and pooling, and KV-cache and data-movement strategies.
Evaluate, benchmark, and select storage and memory technologies against real AI workloads, translating results into clear decisions on performance, capacity, power efficiency, and total cost of ownership.
Operate as the trusted advisor to external customers, sizing solutions, validating requirements, and resolving performance issues in production environments.
Shape the internal platform roadmap by integrating storage and memory into Kubernetes-based and bare-metal environments alongside compute and networking teams.
Bring low-latency fabric expertise (RDMA, RoCEv2, InfiniBand) to connect data systems and establish operational best practices for data protection, resilience, and lifecycle management.
Translate complex data-path decisions into measurable outcomes for latency, throughput, scalability, and total cost of ownership for AI training and inference workloads.
Validate that systems meet the stringent requirements of the AI Factory, ensuring that every watt and every byte are optimised for efficiency and cost.
Act as Firmus's subject-matter expert on data paths across the AI fabric, spanning cache layers, memory fabrics, and high-performance storage systems.
Partner with engineering teams to define data movement strategies that eliminate bottlenecks and maximise GPU utilisation.
Provide thought leadership and technical guidance that influences product direction, integration decisions, and long-term infrastructure strategy.
Requirements
You possess deep, vendor-agnostic expertise across the AI data path, covering storage, memory, and caching layers that keep GPUs fully utilised.
You have hands-on experience with one or more leading high-performance storage and memory platforms in the class of WEKA, VAST Data, DDN, or Dell PowerScale/PowerFlex.
You are fluent in NVMe and NVMe-oF, parallel and software-defined storage, object and file systems, and low-latency network fabrics such as RDMA, RoCEv2, and InfiniBand.
You track next-generation memory directions such as CXL and memory disaggregation, and you understand how these technologies impact AI data infrastructure.
You understand the principles of data protection, resilience, and lifecycle management required by large-scale cloud infrastructure.
You have the ability to translate workload requirements into architecture decisions that balance performance, capacity, and total cost of ownership.
You operate effectively as a trusted advisor, communicating complex technical concepts to both engineering teams and executive stakeholders.
You bring a strong foundation in the fundamentals of data movement, caching, and memory hierarchies that determine real training and inference performance.
You are comfortable making decisions that prioritise efficiency and scalability in a business model built on making every watt and every byte count.