Machine Learning Engineer
Job description
About the role
You will scale up model inference and run predictions at scale on cutting-edge models while owning the end-to-end deployment of machine learning models in production environments. You will work end to end to connect ML models to human interfaces such as APIs, browsers, and applications, ensuring seamless interaction between core systems and user touchpoints. You will design and implement large-scale data and ML pipelines through a full end-to-end product development lifecycle, from initial requirements to stable release. You will collaborate with a cross-functional team of engineers, researchers, product managers, designers, and operations teammates to create cutting-edge products that push technical boundaries. You will bring a high level of end-to-end ownership and self-direction, proactively identifying problems and architecting robust solutions. You will apply your ability to reason through machine learning system tradeoffs to optimize for performance, reliability, and correctness. You will contribute to the broader ecosystem by supporting open-source ML products and infrastructure where applicable.
Key facts
What you'll do
- Scale up model inference and running predictions at scale on cutting-edge models while monitoring system health and performance metrics.
- Work end to end to connect ML models to human interfaces such as APIs, browsers, and applications, translating product requirements into technical solutions.
- Design and implement large-scale data and ML pipelines through a full end-to-end product development lifecycle, including data validation, feature engineering, and model deployment.
- Collaborate with a cross-functional team of engineers, researchers, product managers, designers, and operations teammates to align technical execution with business goals.
- Participate in code reviews, design discussions, and architectural decisions to ensure maintainable and scalable systems.
- Implement experiments to evaluate model performance, analyze results, and iterate based on empirical evidence and user feedback.
- Build and maintain tooling that enables data scientists and product teams to operationalize models efficiently with minimal friction.
- Troubleshoot production issues related to model serving, data quality, and system integration, driving resolutions from detection to prevention.
- Contribute to the definition and evolution of technical standards for data and model workflows across the organization.
- Support the deployment of models in diverse environments, balancing constraints such as latency, throughput, and resource utilization.
- Partner with product teams to prototype new features, validate hypotheses, and deliver measurable impact through data-driven iterations.
- Explore and integrate emerging techniques in machine learning research to maintain a competitive edge in performance and capability.
- Monitor and document system behavior, creating dashboards and reports that provide visibility into model and pipeline health.
- Mentor and guide less experienced team members, fostering a culture of learning, experimentation, and ownership.
Requirements
- Experience as a software engineer with a strong track record of delivering reliable, maintainable software systems in production environments.
- Experience building and serving machine learning models in real-world applications, with evidence of handling complex model behavior in practice.
- Familiarity with Python and related ML frameworks such as PyTorch, Tensorflow, Jax, and other open-source stacks such as HuggingFace, demonstrating fluency in modern ML tooling.
- Ability to reason through machine learning system tradeoffs, including performance, accuracy, scalability, and operational complexity.
- A high level of end-to-end ownership and self-direction, taking initiative to drive projects from conception to completion without excessive supervision.
- Strong problem-solving skills and comfort working with ambiguous, rapidly evolving technical challenges in a startup context.
- Effective communication skills to collaborate with diverse stakeholders and translate technical concepts into actionable insights.
- A commitment to writing clean, tested, and maintainable code that can withstand scrutiny in production systems over time.
- Willingness to work closely with cross-functional partners, including product, design, and operations, to ensure alignment on priorities and outcomes.
- Understanding of software development best practices, including version control, testing, and continuous integration.
Nice to have
- Familiarity with basic machine learning system stacks e.g. TinyML, Triton, CUDA, ROCm, Exo, MLIR, Halide, etc, which can provide deeper insight into system-level optimization.
- Familiarity with DataOps, MLOps, and ML orchestration pipelines, enabling more efficient and reliable workflows across the model lifecycle.
- Understanding of modern ML architectures and intuition for inference performance tradeoffs, helping to balance speed, cost, and accuracy.
- Experience or interest in working on open-source ML products, contributing to community projects and improving shared tooling.
- Interest in building tech aligned with user privacy, computational integrity, and/or censorship resistance, reflecting values around decentralized and secure systems.
- Experience at fast-growing companies or startups, where adaptability, ownership, and rapid learning are essential to success.
Practical notes
The role is fully remote, allowing flexibility in location for the successful candidate. No in-person attendance, travel, or visa requirements are associated with this position at this time. There are no published working hours or application deadlines specified in the source material.