Sr. Computer Vision Engineer
Job description
About the role
You will be an integral part of our innovative team of hardware, full-stack, and computer vision engineers dedicated to advancing retail intelligence solutions. Your primary focus will be on developing, training, and deploying cutting-edge computer vision models that address real-world retail challenges such as inventory management, product recognition, and scene understanding. You will be responsible for designing models that can operate efficiently on edge devices, ensuring high performance and robustness in diverse retail environments. This role requires a versatile engineer who can seamlessly move from model development and fine-tuning to deployment and infrastructure optimization, contributing to the overall strategy of our computer vision capabilities.
Key facts
What you'll do
- Design, develop, and iterate on custom object detection models specifically tailored for retail environments, including inventory tracking, product recognition, and scene analysis. You will work on building models that can accurately identify and classify retail items in various conditions and settings.
- Fine-tune and deploy open-source vision-language models such as LLaVA, Qwen-VL, InternVL, and PaliGemma for enhanced product understanding, zero-shot classification, and scene reasoning. You will build pipelines that integrate these models into real-world perception systems, enabling the system to interpret visual data and translate it into actionable insights.
- Develop vision-language-action pipelines that enable perception systems to reason about scenes and make decisions based on visual input. These pipelines will connect visual understanding with downstream actions, such as inventory updates or alerts.
- Optimize models for edge deployment by applying techniques like quantization, pruning, and architectural modifications to ensure models run efficiently on constrained hardware. You will work with frameworks such as TensorRT and ONNX Runtime to achieve high performance on devices like NVIDIA Jetson.
- Build and maintain robust data pipelines and annotation workflows to improve model accuracy and robustness across diverse retail scenarios. This includes managing datasets, annotations, and data augmentation strategies to enhance model generalization.
- Stay at the forefront of computer vision and vision-language model research by prototyping new architectures and evaluating their suitability for production deployment. You will determine which innovations are ready for real-world use and which are still in the research phase.
- Provide technical leadership by mentoring junior engineers, establishing best practices for model development, and guiding the team on infrastructure and deployment strategies. You will help shape the technical direction of our CV solutions.
- Collaborate with cross-functional teams, including hardware, software, and product teams, to integrate CV models into our retail solutions. You will ensure seamless deployment and operation of models in production environments.
- Participate in code reviews, system design discussions, and documentation efforts to maintain high-quality engineering standards.
- Contribute to the continuous improvement of our ML infrastructure, including version control, CI/CD pipelines, and monitoring systems to ensure reliable and scalable deployment of CV models.
- Engage with stakeholders to understand retail operational needs and translate them into technical solutions that leverage computer vision and multimodal AI.
Requirements
- Minimum of 4 years of hands-on experience in computer vision engineering, with a proven track record of deploying models into production environments.
- Deep expertise with YOLO and YOLO-E architectures, including training, fine-tuning, and understanding their quirks and performance characteristics.
- Practical experience working with open-source vision-language models such as LLaVA, Qwen-VL, InternVL, or PaliGemma. This includes fine-tuning, evaluation, and deploying these models in production settings.
- Familiarity with vision-language-action frameworks and their application to perception and decision-making tasks in real-world scenarios.
- Strong experience with edge deployment frameworks such as TensorRT and ONNX Runtime, including techniques for model quantization, pruning, and optimization for resource-constrained devices.
- Solid software engineering fundamentals, including writing clean, maintainable code, using version control systems like Git, and implementing CI/CD pipelines for machine learning workflows.
- Understanding of the differences between experimental notebooks and production-grade ML systems, with experience in building scalable, reliable systems.
- Ability to troubleshoot and optimize models for performance and accuracy in diverse retail environments.
- Excellent communication skills to collaborate effectively with cross-disciplinary teams and stakeholders.
Nice to have
- Experience developing solutions deployed on NVIDIA Jetson hardware, including Jetson AGX Xavier or Jetson Nano.
- Background working in retail, inventory management, or similar product-focused computer vision applications.
- Proficiency with PyTorch and familiarity with modern training frameworks such as Transformers, LitGPT, or Unsloth.
- Experience running vision-language model inference efficiently using tools like vLLM, llama.cpp, or SGLang.
- Knowledge of synthetic data generation and data augmentation techniques to improve model robustness.
- Familiarity with model versioning and experiment tracking tools such as MLflow or Weights & Biases.
- Contributions to open-source projects or publications related to computer vision or multimodal AI.
- Experience working with AWS cloud services like EC2, ECS, Fargate, S3, Bedrock, or SageMaker for training, deployment, or data storage.
Skills & tools
- PyTorch
- YOLO and YOLO-E architectures
- Open-source vision-language models (LLaVA, Qwen-VL, InternVL, PaliGemma)
- TensorRT and ONNX for model optimization and deployment
- Docker and Kubernetes for containerization and orchestration
- MLOps tools for CI/CD, experiment tracking, and model versioning
- Familiarity with vLLM, llama.cpp, SGLang for efficient inference
- Experience with synthetic data generation and augmentation techniques
Practical notes
This is an on-site role based in the United States, with no remote work option specified. Candidates should be prepared to work closely with hardware teams to optimize models for edge deployment, particularly on NVIDIA Jetson devices. The role involves hands-on model training, deployment, and infrastructure development, requiring a strong understanding of both software engineering and hardware constraints. You will be expected to contribute to the development of scalable, reliable CV solutions for retail environments, ensuring models perform accurately and efficiently in real-world settings. Collaboration across teams and clear communication are essential, as is a proactive approach to staying current with research advancements and integrating innovative solutions into production.