Staff Software Engineer, ML/AI Platform
Job description
About the role
Attentive is seeking an accomplished Staff Software Engineer to join the Machine Learning Platform team as a high-impact individual contributor focused on building the AI and ML infrastructure that powers our AI product suite. You will architect and build the foundational platform components that enable AI and ML engineers and data scientists to train, deploy, and serve models and agentic infrastructure with velocity, performance, and reliability at scale. As a Staff-level IC, you will operate as a technical force multiplier, setting the technical direction for AI and ML infrastructure across Attentive's AI organization and leading through technical excellence. You will advocate for long-term architectural progress while balancing immediate platform needs, executing strategic initiatives measured in quarters and years. Your work will span high-leverage decisions designed to enable entire teams to ship AI and ML capabilities faster and more reliably.
Key facts
What you'll do
Setting Technical Direction
Architect ML platform strategy spanning data pipelines, training infrastructure, and serving layers using cutting-edge tooling like Ray, MLFlow, Metaflow, Argo, and Spark.
Uplevel and Innovate Core AI & ML Stack
Build and operate production-grade, low-latency ML serving layers with robust model lifecycle systems including champion/challenger testing, automated rollouts, versioning, and rollback capabilities.
Uplevel and Innovate Core AI & ML Stack
Define and drive Attentive's agentic stack from the ground up, including infrastructure for MCP, data pipelines, context store, orchestration, and prompt layers.
Technical Leadership
Provide ML infrastructure perspective in high-level discussions about Attentive's AI strategy spanning multiple quarters and teams, influencing priorities and tradeoffs.
Technical Mentorship
Mentor platform and ML engineers, actively championing team members' growth through code reviews, design discussions, and knowledge sharing.
Being the "Glue"
Build universal interfaces, architectures, and patterns - like data access layers and prediction serving APIs - that bridge platform capabilities with product needs to streamline high-priority ML work across the organization.
Scaling Real-Time Workloads
Design and implement inference pipelines with champion/challenger shadow testing and automated model promotion to support continuous validation and safe deployment.
Operational Excellence
Own and evolve core platform components in Kubernetes on AWS, optimizing for performance, reliability, and cost efficiency across batch and real-time workloads.
Enabling Product Teams
Partner closely with product and data science teams to deliver self-service workflows that accelerate experimentation, deployment, and iteration on AI-powered features.
Requirements
You have the experience to know what works, what doesn't, and why in AI and ML systems, allowing you to make informed tradeoffs.
You bring 7+ years focused specifically on ML Platform/MLOps, with deep understanding of gold-standard practices and best-in-class tooling.
You have a proven track record of owning and building core components of ML platforms using tools like Spark, Ray, MLFlow, Kubeflow, or Metaflow.
You have built and operated a high-throughput agentic stack (MCP / data infrastructure, context store, orchestration, and prompt layer) in production environments.
You demonstrate strong expertise in Python for both batch processing and online service frameworks, writing maintainable and scalable code.
You have experience designing and operating online and offline inference systems, understanding the critical differences and tradeoffs between them.
You are comfortable navigating ambiguous problems and driving technical direction in fast-paced, cross-functional settings.
You communicate effectively with both technical and non-technical stakeholders, translating complex infrastructure concepts into actionable plans.
Nice to have
Preferred experience contributing to or building open-source data and ML infrastructure projects.
Deep familiarity with large-scale model training and deployment patterns in cloud-native environments.
Hands-on work with reinforcement learning and agentic workflow orchestration at scale.
Practical notes
This is a full-time position based in the United States.
The compensation range listed applies to US-based applicants and includes salary, equity, and benefits.
By applying for this position, your data will be processed as per Attentive's Privacy Policy.
Attentive is an Equal Opportunity Employer and welcomes applicants from all backgrounds.