Senior/Staff Machine Learning Researcher
Job description
Terra AI at Terraai.
About the role
Probabilistic models and generative methods define subsurface decision-making for critical resource discovery under uncertainty.
Data roles turn raw information into decisions. Analysts query databases and build dashboards. Data scientists build models that predict outcomes. Data engineers build the pipelines that move and store data. All three work closely with business teams and need a mix of statistics, coding, and communication. Nearly every modern company runs on data teams, from startups to banks. A strong portfolio of past analyses matters more than degrees in many hiring decisions.
Key facts
What you'll do
Generative models create 3D geological conditionals that inform exploration decisions for clean energy applications. Each conditional output captures known and unknown subsurface characteristics to guide exploration teams.
Probabilistic modeling fuses geophysical surveys and borehole observations to quantify uncertainty for subsurface targeting. This fusion helps exploration teams reduce risk when evaluating strategic resources.
Synthetic data generation enhances model performance and accelerates discovery timelines for exploration teams. Generated data supports faster, more informed decisions across resource evaluation workflows.
Diffusion approaches adapt to specific projects so that real-world resource constraints guide model behavior and reduce exploration risk. Project teams collaborate to tailor models for specific resource challenges and operational limits.
Requirements
Extensive PyTorch experience is required to write custom modules, optimize training, and debug large-scale models. This expertise supports reliable execution of complex workflows at scale.
Deep learning architecture expertise from scratch is required to design, implement, and train complex models for subsurface problems. Foundational architecture knowledge underpins robust performance in geoscience contexts.
Data curation skills are required to build, clean, and maintain datasets that drive machine learning performance. High-quality datasets directly influence model accuracy and decision confidence.
Strong software engineering and design experience is required, including version control, testing, optimization, and scalable system design. These practices ensure maintainable and reproducible development.
Generative modeling experience with diffusion models and posterior sampling methods is required to condition generation on observations. This experience enables effective integration of physical survey constraints.
Transformer architecture knowledge is required to build and train models for 3D data applications. Transformers help capture spatial and structural patterns in subsurface data.
Scaling expertise across large GPU clusters is required to parallelize models and optimize distributed training pipelines. Distributed training reduces iteration time for model refinement.
Cloud infrastructure experience is required to provision resources, manage costs, and maintain machine learning environments. Stable infrastructure supports continuous training and deployment cycles.
Practical notes
The role operates remotely across the United States with a partnership-focused team. Typical interview steps
Data interviews commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies give a take-home analysis. Expect questions about past projects and the business impact of your work. Interviewers often evaluate how you communicate uncertainty and business impact, not only the math. Bringing a clean write-up of a past analysis to the interview is well received.
Good to know
Machine learning research in geoscience targets subsurface uncertainty reduction for resource exploration. Diffusion models and probabilistic methods support generative modeling for 3D geological conditionals. Large-scale training across GPU clusters relies on cloud infrastructure and distributed optimization. Data curation and software engineering practices underpin reliable model development and deployment.
Career growth
Data careers grow toward senior analyst, staff data scientist, or data engineering lead. Many professionals specialize in machine learning, analytics, or infrastructure. Cross-functional work with product and engineering teams becomes more important at senior levels. The field changes quickly, so continuous learning is part of the job. Professionals who can translate numbers into decisions tend to advance fastest.