Software Engineer, ML Data Infrastructure
Job description
About the role
You will architect and operate the data infrastructure that underpins next generation generative media models at massive scale. This role owns the pipelines that process, transform, and serve the datasets fueling Ideogram's core models. You will collaborate daily with elite research and engineering teams to turn novel data requirements into reliable production systems. The position demands comfort with ambiguity and the ability to drive projects from raw idea to robust, production grade implementation. You will be responsible for ensuring data throughput, reliability, and performance across multi-modal training workflows. This role requires deep collaboration with distributed systems and infrastructure partners to keep the ML engine running smoothly. You will play a key part in shaping the technical roadmap for data storage, processing, and access patterns.
Key facts
What you'll do
Architect and scale data ingestion pipelines to handle petabyte scale datasets across multi-modal sources for foundation model training.
Design and implement robust storage architectures using distributed databases and large scale object stores to support high throughput ML workloads.
Optimize data processing workflows to achieve massive throughput while ensuring reliability, performance, and efficient resource utilization.
Partner with research scientists to translate data requirements into production grade systems that accelerate model development cycles and reduce iteration time.
Build and maintain core infrastructure components on Kubernetes, integrating tightly with GCP services such as Bigtable, BigQuery, Spanner, and Pub/Sub.
Implement monitoring, alerting, and operational tooling to ensure data pipelines are observable, debuggable, and resilient at scale.
Collaborate closely with backend engineers to define data contracts and interfaces that enable efficient model training and evaluation workflows.
Drive projects end to end, scoping technical solutions, executing implementations, and iterating based on feedback and performance metrics.
Leverage distributed systems fundamentals to solve complex scaling challenges across storage, computation, and network resources.
Champion best practices in software engineering, including testing, documentation, and modular design, to ensure long term maintainability.
Work hands on with TPU infrastructure and large scale storage solutions to support the demanding needs of generative media model training.
Identify opportunities for automation and infrastructure improvements, then implement solutions that increase team efficiency and system reliability.
Contribute to technical design discussions, proposing alternatives and trade offs for data architecture decisions that impact scale and performance.
Mentor and guide less experienced engineers by providing code reviews, technical guidance, and knowledge sharing on data infrastructure topics.
Requirements
2 to 5 years developing and shipping large scale distributed systems with proven ability to manage complexity through thoughtful abstractions and scalable design.
Strong fundamentals in data structures, algorithms, and distributed systems that underpin high performance data infrastructure.
Deep understanding of databases and data storage architectures, including relational and non relational models at scale.
Hands on experience with large scale data processing systems, including stream and batch processing patterns in production environments.
Demonstrated ability to drive projects from 0 to 1, including scoping, execution, iteration, and delivery against demanding timelines.
A deep sense of ownership that drives you to identify opportunities, suggest improvements, and act decisively on impactful changes.
Thrives in fast moving, ambiguous environments with a strong bias toward action and a results oriented mindset.
Asks incisive questions, thinks from first principles, and seeks out resources to continuously deepen technical understanding.
Comfortable working with modern backend tools, including container orchestration with Kubernetes and infrastructure as code using Terraform.
Experience with GCP managed services such as Google Bigtable, Google BigQuery, Google Spanner, and Google Pub/Sub is essential.
Strong written and verbal communication skills for effective collaboration with cross functional research and engineering teams.
Nice to have
Experience with TPU infrastructure and large scale machine learning training pipelines.
Background in generative media, graphics, or related AI applications.
Contributions to open source data infrastructure or distributed systems projects.
Practical notes
Work is full time based in Toronto, with access to the NYC office as needed.
The role operates standard business hours with flexibility for deep work and collaboration across time zones.
Ideogram provides comprehensive support for visa and immigration requirements for eligible candidates.
Competitive compensation and equity are offered to recognize the impact and value of contributions to company success.