
Senior Machine Learning Engineer I, Physical Sciences
Job description
About the role
This position centers on owning the full lifecycle of scalable machine learning workflows for the physical sciences. You will guide projects from initial intake to final delivery, advancing research that contributes to scientific superintelligence. The role demands deep technical ownership and a commitment to solving high-impact problems in materials, chemistry, and physics. You will architect and implement solutions that translate complex scientific questions into robust, automated pipelines. A significant portion of your time will be dedicated to ensuring that models generalize well beyond narrow datasets. You will be responsible for maintaining the integrity and reproducibility of experiments across the entire machine learning lifecycle. This role requires a proactive approach to identifying bottlenecks and designing systems that empower researchers to iterate quickly.
Location, Engagement, and Compensation
This is a full-time position based in Cambridge, MA, USA. The expected base salary range is $148,000 to $198,000 USD, complemented by bonus potential and early-stage equity. Final compensation will be determined based on background, expertise, and anticipated impact.
What you'll do
- Architect end-to-end machine learning pipelines specifically designed for the demanding constraints of physical sciences research.
- Build and deploy production-grade data ingestion frameworks capable of handling heterogeneous and noisy scientific datasets.
- Design feature engineering strategies that capture the intrinsic physical relationships within complex materials and chemical structures.
- Implement model training regimes that leverage high-performance computing resources to accelerate the development of predictive models.
- Establish rigorous evaluation methodologies to assess model performance against scientific validity and not just statistical accuracy.
- Deploy scalable serving infrastructures that deliver low-latency predictions for real-time research and discovery workflows.
- Implement comprehensive monitoring systems to detect data drift and model degradation in long-running scientific experiments.
- Define and maintain detailed documentation that ensures reproducibility and facilitates knowledge transfer across interdisciplinary teams.
- Lead technical design reviews and establish coding standards that promote maintainable and efficient scientific software.
- Integrate closely with infrastructure teams to optimize resource utilization and manage workflow orchestration at scale.
- Take ownership of Large Language Models, multimodal architectures, and RAG systems to extract insights from scientific literature.
- Profile and debug complex model pipelines to resolve performance bottlenecks and ensure system reliability.
- Optimize GPU utilization through advanced techniques involving CUDA, Triton, and compilation strategies for maximum throughput.
- Contribute code and enhancements back to open-source ML and scientific software communities.
- Ensure seamless integration with existing large-scale compute environments used by research scientists.
- Manage data provenance to provide complete visibility into the lineage of training data and model outputs.
- Collaborate with domain experts to translate scientific hypotheses into testable machine learning experiments.
- Mentor junior engineers and promote best practices to elevate the overall capability of the machine learning organization.
Requirements
- Hold a Bachelor of Science, Master of Science, or Doctor of Philosophy in Computer Science, Engineering, or another related quantitative field.
- Possess equivalent industry experience that demonstrates mastery of machine learning engineering principles.
- Exhibit a strong foundation in Python software engineering with expertise in testing, packaging, and type hinting.
- Demonstrate extensive experience with machine learning frameworks, specifically PyTorch and Huggingface, in production environments.
- Have a proven track record of deploying ML services into cloud-based infrastructure reliably and efficiently.
- Show proficiency with FastAPI, GRPC, containerization technologies, and orchestration tools such as Kubernetes.
- Provide evidence of deploying models in production settings, particularly involving LLMs, multimodal models, databases, and RAG systems.
- Maintain sharp debugging and profiling skills to diagnose issues in complex distributed systems.
- Communicate and collaborate effectively at the intersection of scientific inquiry and engineering execution.
- Thrive in an environment where scientific curiosity drives the development of sophisticated algorithms.
- Adhere to strict standards for code quality, reliability, and operational excellence.
- Understand the importance of data privacy and security when handling sensitive research information.
- Be comfortable working autonomously while also contributing to a tightly integrated team environment.
- Commit to continuous learning to keep pace with rapid advancements in machine learning and scientific computing.
- Apply problem-solving skills to ambiguous challenges common in early-stage scientific research.
- Take responsibility for the long-term maintainability of the codebase and infrastructure.
- Align with the mission of advancing scientific discovery through the application of machine learning.
Nice to have
- Exposure to scientific or engineering domains such as materials, chemistry, or physics, along with familiarity with related data formats or benchmarks.
- Hands-on experience with GPU optimization techniques, including CUDA, Triton, and distributed training frameworks.
- A history of contributions to open-source machine learning or scientific software ecosystems.
- Background with workflow orchestration, data provenance, and large-scale compute environments common in scientific research.
Practical notes
No specific tools are listed for this role. All candidates are encouraged to apply.