GRASP-CONDITIONED DEPENDENCY GRAPHS

Minimal-Intervention Target Retrieval
in Cluttered Scenes

Which objects block this grasp?

Wanhao Niu1,*, Qiyan Ke1,*, Yuan Sun2, Chongrui Zhang3,
Huaidong Zhou1,†, Yongfeng Rong1, Shuang Zhang4, Fuchun Sun1,†

1 Department of Computer Science and Technology, Tsinghua University
2 School of Reliability and Systems Engineering, Beihang University
3 Department of Mechanical Engineering, Tsinghua University
4 School of Engineering, Hebei Normal University

* Equal contribution   † Corresponding authors

GCDG overview: candidate grasps, object–grasp dependency prediction, and minimal-intervention target retrieval
The same surrounding object can block one grasp while leaving another feasible. GCDG makes this dependence explicit.

Reason about actions,
then remove only what matters.

Minimal-intervention retrieval asks a robot to extract a target from clutter while limiting non-target removals. GCDG represents candidate grasps as action nodes and predicts typed dependencies on object–grasp edges. The planner compares candidate grasps through their blockers and updates its plan after every intervention.

1

Intervene

Fix a target grasp. Remove one surrounding object and test whether approach, lift or full feasibility improves.

2

Predict

Use observed object features, grasp geometry and relational context to predict four dependency labels per edge.

3

Retrieve

Choose a grasp and a short blocker-removal prefix. Refresh the observation and replan after the scene changes.

GCDG DENSE BENCHMARK · V1.0

Dense scenes.
Grasp-specific supervision.

A benchmark for dependency prediction and minimal-intervention planning, with 25–27 objects per dense scene, parallel-jaw and suction proposals, typed intervention labels and minimal blocker sets.

500dense scenes
12,985target samples
313,228grasp candidates
6.77Mlabeled edges
Fixed scene-disjoint splits
SplitScenesTarget samplesParallel-jawSuction
Train3509,090112,521106,428
Validation751,95024,46222,792
Test751,94524,32122,704

The annotation release includes fixed splits and SHA-256 manifests. Stored graph inputs can be used without running a simulator or a native grasp SDK. Third-party meshes, textures and native proposal backends are acquired separately.

python scripts/download_benchmark.py --split all --output data --extract

From dependency prediction
to physical retrieval.

Evaluate the representation, the planning trade-off and the execution interface separately. Each task has a fixed cohort and protocol.

75 test scenes · 1,945 target samples · mean ± sample SD over seeds 7/11/23

MethodAP (%) ↑F1 (%) ↑Blocker IoU (%) ↑

G2N2 is adapted to the common dependency-prediction interface. Full ablations and all seeds are included in the release.

Explore all 20 tables and underlying results →
Real-robot target-retrieval experiment from Figure 5 of the camera-ready paper
Real-robot retrieval with the shared grasping interface. Hardware results use three recorded batches of twenty trials per method.

Start with one graph.

Inspect the included real benchmark sample, reproduce saved-result counts, or download the full dataset to train and evaluate the released architectures.

git clone https://github.com/wanhaoniu/GCDG.git
cd GCDG
python -m pip install -e .
python scripts/inspect_sample.py --dataset examples/benchmark
python scripts/recount_ordered.py
Full reproduction guide →

Citation

@misc{niu2026gcdg,
  title={Minimal-Intervention Target Retrieval in Cluttered Scenes
         with Grasp-Conditioned Dependency Graphs},
  author={Niu, Wanhao and Ke, Qiyan and Sun, Yuan and Zhang, Chongrui
          and Zhou, Huaidong and Rong, Yongfeng and Zhang, Shuang
          and Sun, Fuchun},
  year={2026},
  url={https://github.com/wanhaoniu/GCDG}
}

Acknowledgments

Supported by the Beijing Natural Science Foundation (Nos. L233006 and L253006), and the Joint Funds of the National Natural Science Foundation of China (No. U22A2057).

Correspondence: Huaidong Zhou and Fuchun Sun.