On-device Embodied World Models Workshop @ ECCV 2026
Autonomous robotic grasping in highly cluttered environments, such as industrial bins, warehouse logistics totes, or chaotic domestic spaces, remains a fundamental bottleneck in robotics. While modern vision systems excel at identifying isolated, clearly visible surface-level objects, achieving true operational autonomy requires models to process complex semantic and geometric physical hierarchies.
To grasp a buried or partially obscured item safely, an intelligent agent cannot merely rely on standard vision-language grounding. Instead, it must implicitly map the environment, reason systematically about obstructions and physical occlusions, and determine the exact sequence of clearing actions required to safely unlock an unobstructed trajectory to the target object.
The core objective of the UNOBench Challenge is to accelerate the development of deep learning frameworks, specifically Vision-Language Models (VLMs) and multi-modal embodied network architectures capable of executing visually-grounded obstruction reasoning.
Participants will train and evaluate their approaches using the newly curated UNOBench dataset. Built on top of MetaGraspNetV2, UNOBench provides a rigorous benchmark for obstruction reasoning, in scenarios with different complexity levels.
Winners of the challenge will be able to present their work at the On-device Embodied World Models (ODEWM) Workshop at ECCV 2026.
Submissions will be evaluated and ranked based on their ability to identify all the objects that should be removed first to grasp a particular object (i.e., all the last object in the occluding chain).
The challenge offers two main track:
All deadlines are strictly enforced at 23:59 UTC on the dates listed below:
| Phase / Milestone | Date | Details |
|---|---|---|
| Challenge & Data Launch | June 13, 2026 | UNOBench challenge data, and baseline model (UNOGrasp) code released. |
| Evaluation Phase | July 1st, 2026 | The evaluation server opens. Live public leaderboard tracks participant validation submissions. |
| Final Test Submissions | August 30, 2026 | Final submission deadline. |
| Winner Announcement | September 1, 2026 | Top performing architectures finalized following verification against cheating/overfitting. |
| Workshop Day | September 8, 2026 | Winning teams present technical talks alongside invited keynotes at our official conference venue. |
Two tracks, one winner. Choose your path.
Test your foundational knowledge in this introductory challenge.
A much harder test for those who have mastered the basics.
Organizing Secretariat & Contact: For inquiries regarding
automated evaluation configurations, API usage constraints, or academic
dataset licensing, you can use our mailing list
unobenchchallenge@fbk.eu or join our google group,
please note that all messages using these channels are visible to all members of the mailing list.
For private inquiries, you can contact Francesco Giuliari at fgiuliari@fbk.eu
This challenge has been organized by the Technologies of Vision unit at Fondazione Bruno Kessler.
The UNOBench challenge and dataset are based on the work presented at CVPR 2026:
@InProceedings{Jiao_2026_CVPR,
author = {Jiao, Runyu and Bortolon, Matteo and Giuliari, Francesco and Fasoli, Alice and Povoli, Sergio and Mei, Guofeng and Wang, Yiming and Poiesi, Fabio},
title = {Obstruction Reasoning for Robotic Grasping},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {20755-20764}
}
This site uses cookies from Google Analytics to help analyze traffic and improve your experience. No personal data is collected. You can choose to accept or reject these analytics cookies.