Software Engineer, ML Platform
Job description
, Inc..
About the role
Gusto is seeking a strong Machine Learning Platform and Infrastructure Engineer to join our ML Platform team and build out and scale our ML and AI platform. In this role, you will own the design, implementation, and reliability of the core infrastructure that powers machine learning workflows across the company. You will partner closely with AI and ML engineers to translate model development needs into scalable platform capabilities. Your work will ensure that models can be built, deployed, monitored, and retrained with consistency and efficiency at every scale. You will also play a key role in defining and documenting optimal processes that set the standard for how ML infrastructure is built and maintained.
Key facts
What you'll do
Build core components of our ML and AI Platform technical roadmap to design and build MLOps solutions with automated pipelines and standardized processes to build, deploy, run, monitor, debug, and retrain ML and AI Models.
Develop, maintain, and enhance frameworks for machine learning model development and deployment.
Collaborate with the ML/AI builders and application owners to determine business requirements and SLAs for API-enabled services.
Develop, maintain, and enhance infrastructure supporting machine learning services.
Support the development of new patterns for the deployment of machine learning models with CI/CD pipelines and automated testing.
Apply AI tools as a regular part of your engineering workflow, and bring an AI-native lens to engineering and product decisions: identify where AI can reduce effort, simplify complex workflows, and surface proactive guidance.
Adopt the latest best practices for using AI technologies across all aspects of technical development.
Work closely with data teams to ensure that data pipelines feeding ML models are robust, scalable, and observable.
Implement monitoring and alerting for model performance, data drift, and infrastructure health to enable rapid troubleshooting and iteration.
Partner with security and compliance stakeholders to ensure that ML infrastructure and models adhere to governance standards.
Design experiments to validate infrastructure changes and measure impact on model performance and developer productivity.
Document system architecture, operational runbooks, and onboarding guides to support long-term maintainability and team growth.
Contribute to open source tools and internal libraries where appropriate to accelerate development and share best practices.
Engage in code reviews and pair programming to maintain high standards of quality and knowledge sharing across the ML Platform team.
Mentor junior engineers by providing clear technical guidance and feedback on design and implementation decisions.
Requirements
At least 5+ years of software engineering experience with Python, Ruby, or Java.
Demonstrated experience designing and developing infrastructure and platform services for the machine learning lifecycle, such as feature stores, model development, deployment, and observability tools and solutions.
Experience with at least one of the major cloud platforms, with AWS preferred but not required.
Curiosity and experimentation with emerging AI frameworks, applying and sharing best practices to evaluate and scale AI use safely across teams.
Comfort with AI-assisted development tools and a habit of staying current with emerging approaches to building software.
Strong understanding of machine learning concepts, including training workflows, inference patterns, and common model evaluation metrics.
Ability to write clean, maintainable, and well-tested code that scales efficiently under load.
Experience with infrastructure-as-code practices and tools for provisioning and managing cloud resources.
Nice to have
Experience contributing to or using open source ML infrastructure projects.
Familiarity with Kubernetes and container orchestration for ML workloads.
Knowledge of feature store implementations and versioning strategies.
Experience with distributed training patterns and model serving frameworks.
Practical notes
This is a full-time position based in the United States. Compensation varies by location within the stated ranges. Employment eligibility requirements apply.