A tactile-driven framework that grasps an exact, user-specified number of objects from dense clutter — verifying the in-hand count by touch and regrasping to correct itself.
We investigate the novel task of quantitative robotic grasping and propose QTac, a tactile-driven framework designed to grasp a user-specified number of objects. Existing vision-only approaches struggle with severe occlusions in such tasks, making tactile feedback essential for reliably verifying the in-hand quantity. However, applying this concept faces two key challenges: the inherent unreliability of single-attempt grasping paradigms, and the difficulty of integrating high-dimensional tactile signals into existing manipulation models.
QTac addresses these challenges through two key innovations: (1) a human-inspired demonstration strategy that endows the robot with iterative regrasping capabilities to rectify initial count mismatches, reducing the reliance on complex sensory feedback hardware in teleoperation; (2) a decoupled tactile perception module that translates high-dimensional tactile signals into explicit quantitative feedback, directly guiding corrective regrasping. Comprehensive real-world experiments across diverse objects validate that QTac achieves reliable quantitative grasping in dense clutter and successfully extends its quantification capability to novel objects.
Direct-grasping trajectories give dense contact-frame supervision for CountNet, while free-adjustment trajectories provide the regrasp-rich primitives that teach the policy to correct count mismatches — without extra sensory hardware.
A decoupled module maps bilateral tactile depth maps to a probabilistic count distribution. Keeping perception explicit and standalone lets it transfer to novel objects without retraining.
A conditional flow-matching policy fuses the count distribution, target quantity, vision, and proprioception to synthesize corrective actions, iteratively regrasping until the achieved count matches the target.
The policy continuously performs exploratory adjustments until the explicitly decoded tactile state matches the target count. Successful quantitative grasps prior to placement are highlighted; * denotes novel objects.
| Task | Vision-Only | End-to-End | PC [42] | π0.5 (V) [43] | π0.5 (V+T) | Ours (Gen.) | Ours (Spec.) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MAE | SR | MAE | SR | MAE | SR | MAE | SR | MAE | SR | MAE | SR | MAE | SR | |
| T1 · Envelope | 0.87 | 33.3 | 0.67 | 43.3 | 1.07 | 30.0 | 1.53 | 16.7 | 2.27 | 6.7 | 0.53 | 50.0 | 0.43 | 56.7 |
| T2 · Chip | 1.20 | 23.3 | 1.17 | 23.3 | 1.43 | 26.7 | 2.00 | 10.0 | 1.90 | 13.3 | 0.77 | 36.7 | 0.50 | 50.0 |
| T3 · Pencil | 0.93 | 23.3 | 0.83 | 26.7 | 1.10 | 23.3 | 0.87 | 23.3 | 1.33 | 10.0 | 0.70 | 33.3 | 0.43 | 60.0 |
| T4 · Hook | 0.70 | 43.3 | 0.50 | 53.3 | 0.63 | 26.7 | 0.80 | 33.3 | 0.60 | 53.3 | 0.40 | 60.0 | 0.23 | 76.7 |
| Mean (T1–T4) | 0.93 | 30.8 | 0.79 | 36.7 | 1.06 | 26.7 | 1.30 | 20.8 | 1.53 | 20.8 | 0.60 | 45.0 | 0.40 | 60.9 |
| T5 · Ring * | 0.87 | 40.0 | 1.07 | 36.7 | 1.80 | 13.3 | 1.47 | 20.0 | 1.27 | 20.0 | 0.63 | 56.7 | N/A | N/A |
| T6 · Rod * | 0.50 | 53.3 | 0.73 | 40.0 | 0.70 | 36.7 | 0.73 | 36.7 | 0.67 | 40.0 | 0.33 | 66.7 | N/A | N/A |
| Mean (T1–T6) | 0.85 | 36.1 | 0.83 | 37.2 | 1.12 | 26.1 | 1.23 | 23.3 | 1.34 | 23.9 | 0.56 | 50.6 | – | – |