Agentic AI/ML Engineer, Multimodal
Job description
About the role
You will architect and ship core components of Field AI's Field Foundation Model, owning the full lifecycle from raw multimodal data to deployed agentic capabilities in production environments. This role centers on building risk-aware systems that leverage real robot data to improve perception, memory, and tool use across a globally distributed fleet. You will design and implement multimodal retrieval-augmented generation pipelines that enable long-memory video analysis and precise scene understanding in dynamic, unstructured settings. A critical part of your work will involve turning research ideas into robust software, integrating vision, language, and action models that operate reliably on hardware in the field. You will collaborate closely with robotics engineers and field teams to ensure that models are not only accurate but also safe, efficient, and aligned with real deployment constraints. Your contributions will directly shape how Field AI transforms streaming multimodal observations into actionable insights for customers and internal engineering teams. You will own experimental design, rapid prototyping, and rigorous evaluation, iterating on models based on live feedback from deployed systems. This position requires a deep commitment to reliability, observability, and continuous improvement as data and environments evolve over time.
Key facts
What you'll do
Design and implement end-to-end data pipelines that ingest, clean, and curate massive multimodal streams from autonomous robots operating in diverse real-world conditions.
Conduct applied research and model development in computer vision, vision-language models, and multimodal scene understanding, with a focus on agentic AI capabilities such as tool use and memory.
Build and maintain multimodal retrieval-augmented generation systems that support long-memory video analysis, enabling efficient search and reasoning over historical robot experiences.
Optimize model inference for production environments, balancing accuracy, latency, and resource utilization on edge and cloud infrastructures.
Collaborate with field operations to collect, label, and validate data from real deployments, ensuring that models generalize across environments and use cases.
Partner with robotics engineers to integrate perception and insight models directly onto hardware, supporting real-time decision-making and adaptability.
Create and track detailed benchmarks and evaluation frameworks that measure not only accuracy but also robustness, safety, and operational efficiency in deployed settings.
Contribute to the FiFM roadmap by prioritizing features, managing experiments, and documenting model behavior, limitations, and evolution over time.
Participate in cross-functional design discussions that shape the broader perception and insight-generation architecture across Field AI products.
Drive continuous improvement by analyzing failure modes, conducting root-cause analysis, and implementing fixes that improve model performance in the wild.
Maintain strong scientific rigor through experimentation, code reviews, and reproducibility, ensuring that results are measurable and actionable.
Communicate technical findings and trade-offs clearly to both technical and non-technical stakeholders, aligning on goals and success criteria.
Take ownership of model artifacts, monitoring dashboards, and deployment pipelines, ensuring that systems remain reliable and scalable as the fleet grows.
Support on-call responsibilities as needed, responding to model regressions, data quality issues, and deployment incidents in a timely manner.
Requirements
Demonstrated experience as an AI/ML Engineer or researcher building and deploying models in production environments, with a strong track record of shipping reliable systems.
Strong proficiency in Python and modern deep learning frameworks, with hands-on experience building and training neural networks for perception or language tasks.
Solid understanding of computer vision fundamentals, including image classification, object detection, segmentation, and emerging multimodal vision-language models.
Experience with data pipeline construction, dataset curation, and data quality assurance for large-scale, real-world multimodal datasets.
Proven ability to work with vision-language models, including fine-tuning, prompt engineering, and integration into agentic workflows that involve tool use and memory.
Experience designing and implementing retrieval-augmented generation systems, especially for long-context and long-memory applications over video and sensor streams.
Comfortable working in fast-paced, ambiguous problem spaces, where requirements evolve based on field data, hardware constraints, and customer needs.
Strong engineering habits, including modular code design, testing, debugging, and documentation that enable collaboration across distributed teams.
Nice to have
Experience with autonomous robot systems, real-time perception, and edge deployment is highly valued.
Background in building and operating globally deployed AI systems that improve through continuous field data.
Knowledge of robotics simulation tools, sensor drivers, and middleware commonly used in robotic deployments.
Familiarity with MLOps practices, experiment tracking, and model monitoring in production environments.
Experience contributing to open source ML libraries or publishing research in top-tier AI/ML venues.
Practical notes
This is a full-time position based in Irvine, California.
The role may require travel to field deployment sites for hands-on testing and data collection as needed.
Candidates must be authorized to work in the United States without sponsorship for this role.