职位描述
About Caper
Caper (跃氪) is a pioneering company in retail automation technology under Instacart, a North American NASDAQ-listed company. As its core smart hardware brand, Caper is committed to redefining the shopping experience through groundbreaking product design and engineering innovation.
Our flagship product, the Caper Smart Cart, stands as a benchmark for embodied AI technology applications. To achieve seamless and reliable autonomous shopping functionality, we integrate cutting-edge AI vision systems, high-precision sensors, and advanced visual-language model (VLM) algorithms.
The Caper Smart Cart has successfully served both North American and global markets, including dozens of top-tier retail brands such as the top 3 supermarket giants in North America, demonstrating the excellence of our technology and product design. Backed by our parent company Instacart—a global leader in online fresh grocery delivery with over a decade of experience serving 600+ retail brands and 55,000+ stores across North America—Caper benefits from unparalleled market insights, technical resources, and large-scale implementation platforms, providing engineers with vast opportunities to tackle real-world complex challenges.
Welcome to visit our official website: www.caper.ai
—--------------------------------------
About the Job
About the Team
The Spatial Intelligence team builds the core AI technologies that enable smart shopping carts to understand and interact with retail environments. Our mission is to create a persistent Retail World Model that powers real-time cart localization, item localization, shopping assistance, and future autonomous retail robotics. We combine computer vision, localization, semantic understanding, and AI to build a spatial representation of the physical world that continuously evolves as the store changes.
Responsibilities
• Design and develop real-time perception systems for smart shopping carts operating in retail environments.
• Build AI models for shelf understanding, product detection, barcode recognition, OCR, aisle recognition, and semantic scene understanding.
• Develop robust semantic observations that combine recognition results with geometric information and confidence estimates.
• Build multi-camera perception pipelines for dynamic retail environments.
• Develop embedding-based and foundation-model-based approaches for open-world product and scene understanding.
• Optimize AI models for deployment on NVIDIA Jetson platforms using TensorRT, CUDA, FP16, and other inference optimization techniques.
• Design scalable perception architectures supporting multiple AI models, asynchronous inference, and efficient GPU utilization.
• Develop data-centric AI workflows including dataset design, active learning, evaluation, and continuous model improvement.
• Collaborate closely with the Localization team to improve localization robustness through semantic observations.
• Collaborate with the Spatial Modeling team to define semantic observation interfaces that support persistent item localization and world modeling.
• Evaluate emerging research in computer vision, foundation models, multimodal AI, and edge AI for future product capabilities.
About You
Minimum Qualifications
• MS or PhD in Computer Science, Computer Vision, Robotics, Electrical Engineering, or a related field (or equivalent industry experience).
• 5+ years developing production computer vision or AI systems.
• Strong experience with modern deep learning frameworks such as PyTorch.
• Experience with object detection, segmentation, OCR, visual embeddings, or multi-object tracking.
• Multi-view geometry
• Experience deploying deep learning models on GPU edge platforms.
• Strong software engineering skills in Python and C++.
• Experience designing production-quality AI software.
Preferred Qualifications
• Experience with Vision Transformers, CLIP, DINOv2, Grounding DINO, Segment Anything, or other vision foundation models.
• Experience with multi-camera perception or 3D computer vision.
• Experience with TensorRT, CUDA, ONNX Runtime, or GPU performance optimization.
• Experience in robotics, autonomous systems, or edge AI.
• Experience building large-scale production ML systems.
• Publications or open-source contributions in computer vision or AI.
—--------------------------------------
We offer a competitive salary package
At Caper, the total compensation consists of base salary and equity in the form of RSUs (Restricted Stock Units). The job description shows only the base salary component, but we will additionally offer additional stock rewards from our parent company, Instacart. As a publicly traded company in the U.S., Instacart's stock can be directly traded on the market and holds potential for appreciation.
Caper has our own salary philosophy. The final salary of an offer will have a direct connection with the candidate's experience and skill level presented during the interview.
What benefits we aim to provide:
New employees can enjoy 10 days of paid annual leave, 10 days of paid sick leave, and male employees will have an additional 4 weeks of paid paternity leave.
Annual learning fund, annual personal living fund, annual team building activities, health & medical coverage, and more.
Who should join us!
We are looking for driven individuals who thrive in a fast-paced engineering/product environment, are passionate about product and improving team performance, and feel comfortable pushing the limits of what is possible.
We welcome every individual who is direct and open to facilitate our adorable no-political culture.
Every day we solve incredibly complex problems to create an experience for our users that is absolutely magical. Join us!