Senior Software Engineer, ML/AI Platform
Job description
About the role
Attentive is seeking a Senior Software Engineer to join our Machine Learning Platform team. You will build and maintain the infrastructure that allows our data scientists and ML engineers to train, deploy, and monitor models at scale. In this capacity, you will own the design and execution of core platform components that abstract complexity and deliver reliable, self-service capabilities. You will be responsible for ensuring that the end-to-end ML lifecycle is performant, observable, and secure across distributed environments. This role requires a deep understanding of both software engineering best practices and the unique demands of production machine learning workflows. You will partner closely with cross-functional stakeholders to translate product requirements into robust technical solutions. Your work will directly influence the speed and quality with which our organization delivers AI-powered features to market.
Key facts
What you'll do
- Create and maintain production services for the full ML lifecycle, including model training, evaluation, and serving.
- Develop APIs and self-service tools to improve developer productivity for ML workflows.
- Manage projects from initial design through to production monitoring and support.
- Improve the performance, cost efficiency, and reliability of platform components.
- Collaborate with infrastructure, product engineering, and data science teams to deliver AI initiatives.
- Mentor team members and establish high standards for documentation and code quality.
- Implement scalable data pipelines that integrate diverse data sources into reliable feature stores.
- Design and operate infrastructure components that support distributed training and real-time inference.
- Drive standardization across the ML stack by codifying patterns and automating repetitive platform tasks.
- Conduct in-depth analyses of platform incidents to identify root causes and implement preventative measures.
Requirements
- 5+ years of experience in distributed systems or production software development, with a focus on MLOps or ML infrastructure.
- Proficiency in Java, Python, or a similar language.
- Experience with at least one part of the ML lifecycle, such as feature platforms, orchestration, or model deployment.
- Familiarity with cloud environments using AWS, Kubernetes, and infrastructure-as-code.
- Experience working with distributed compute or data systems like Kafka, Ray, or Spark.
- Ability to troubleshoot scalability and reliability issues across infrastructure and application layers.
- Proven track record of delivering complex technical projects independently.
- Strong understanding of software development principles including version control, testing, and continuous integration.
Skills & tools
- Infrastructure: AWS, EKS, Kubernetes, Terraform, Helm, Istio, Datadog
- ML Stack: Metaflow, MLflow, Argo, PyTorch, TensorFlow, Hugging Face
- Data Systems: DynamoDB, Postgres, Redis, Kinesis, Spark, Ray, Kafka
- Languages: Python, Java
Practical notes
- Compensation includes base salary, equity, and benefits.
- Attentive is an equal opportunity employer committed to a diverse and inclusive workplace.
- Reasonable accommodations are available for candidates with disabilities during the hiring process.
- Applicant data is processed according to the Attentive Privacy Policy.
This position is based in the United States and operates on a full-time engagement basis. The successful candidate will work closely with the ML/AI Platform team to advance the company's technical capabilities. You are expected to operate with a high degree of autonomy while adhering to strict standards for code quality and system reliability. The role involves significant responsibility for both architectural decisions and day-to-day execution. You will be evaluated on your ability to deliver measurable improvements to platform stability, scalability, and developer experience. The work you perform will support critical business functions and directly impact the company's strategic objectives. Effective communication and collaboration are essential, as you will interface with multiple teams across the organization. This role is not eligible for relocation assistance. Applicants must be authorized to work in the United States. The company reserves the right to update tooling and technology stacks based on evolving business needs. You should be comfortable working in a fast-paced environment where priorities may shift based on market demands. Regular participation in on-call rotations may be required to ensure platform uptime. Documentation of processes and decisions is a key component of the role, and you are expected to maintain clear and up-to-date records. Training and growth opportunities will be provided to help you keep pace with the latest developments in machine learning infrastructure. This position requires a deep commitment to engineering excellence and a passion for building systems that empower other technologists.