Staff Software Engineer, MDLC
Job description
About the role
At Domino, you will own the design and delivery of the core platform capabilities that power the Model Development Lifecycle for the world's largest AI-driven organizations. You will architect and build the foundational systems that enable seamless experimentation, training, and deployment of machine learning models at enterprise scale. You will collaborate closely with data scientists, product managers, and infrastructure teams to translate complex requirements into robust, scalable software solutions. You will be responsible for extending Domino's Extensions framework, enhancing multi-agent workflow orchestration, and ensuring platform reliability and performance. You will play a key role in shaping how customers operationalize large language models and manage model artifacts across regulated industries. Your work will directly impact how Johnson & Johnson, GSK, UBS, and the US Navy solve critical challenges in drug discovery, finance, and national security. You will thrive in a startup-like environment where ownership, speed, and technical excellence are expected every day.
Key facts
What you'll do
Implement and enhance platform features that enable teams to design, test, and deploy multi-agent workflows at scale across diverse cloud environments.
Expand the platform's inference infrastructure to support high-throughput, low-latency serving of large language models, helping customers confidently operationalize LLM applications at enterprise scale.
Design and integrate secure, scalable APIs such as RESTful services and gRPC endpoints to ensure consistent model delivery on-premises, in enterprise infrastructures, or through third-party hosting.
Collaborate with cross-functional teams to integrate back-end systems with front-end interfaces and third-party services, ensuring smooth data flow and feature cohesion.
Leverage advanced tools like GPUs, Ray, and Spark to build scalable training resources that meet the needs of complex and diverse AI projects.
Use robust testing frameworks including unit, integration, and end-toend tests, and establish CI/CD pipelines that automate validation and deployment.
Profile and optimize back-end performance, focusing on cloud environments and container technologies such as Docker and Kubernetes.
Utilize the Domino model registry to version, store, and enable easy discovery of models across the organization, improving collaboration and reproducibility.
Support organizations in developing, registering, and scaling AI models to enable impactful insights and innovation across the enterprise.
Maintain and evolve core platform components with a focus on reliability, scalability, and security for regulated customers.
Contribute to architectural decisions that balance innovation with operational stability and long-term platform scalability.
Participate in on-call rotations and incident response to ensure platform availability and rapid resolution of production issues.
Requirements
You have a Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
You have hands-on experience developing and managing high-performance back-end systems in distributed computing environments.
You have demonstrated experience building, deploying, and operating scalable APIs, including RESTful APIs and gRPC.
You are proficient in back-end development languages such as Python, Java, Scala, or Go.
You have a strong understanding of performance optimization techniques, especially in cloud environments and with container technologies like Docker and Kubernetes.
You have experience with CI/CD practices and robust testing frameworks such as unit, integration, and end-to-end testing.
You are familiar with traditional machine learning model development and AI workflows, including experiment tracking, hyperparameter optimization, model evaluation frameworks, and managing model artifacts.
You have worked with distributed computing frameworks such as Apache Spark, Azure ML, or SageMaker, and with cloud platforms such as AWS, Azure, or GCP.
Nice to have
Experience with Apache Spark, Azure ML, or SageMaker is a plus.
Proficiency with cloud providers such as AWS, Azure, or GCP and deploying services in these environments.
Hands-on background in building and operating high-throughput, low-latency serving layers for large language models.
Experience extending platform frameworks to enable customer-specific modules and integrations.
Familiarity with model registry systems, versioning, and artifact management for regulated environments.
Experience with collaborative and reproducible machine learning workflows.
Practical notes
This is a full-time role.
Candidates must be eligible to work in the United States.
No visa sponsorship is mentioned in the source text.
There are no application deadlines specified in the source text.