Enhancing AI Binocular Vision for Camellia oleifera Fruit Localization Using Zero-Shot Lightweight Segmentation Model and 3D Surface Optimization
Shouxiang Jin , Lei Zhou , Linyun Xu , Minghong Shi , Hongping Zhou
Engineering ›› : 202606020
The segmentation and three-dimensional (3D) localization accuracy of Camellia oleifera fruits by harvesting robots are critical factors influencing their success rate. To address the challenges of high annotation costs and significant localization errors in fruit segmentation and localization, an artificial intelligence (AI)-based binocular vision method for Camellia oleifera fruit localization is proposed using zero-shot lightweight segmentation and 3D surface optimization models. First, the large vision model Grounding DINO generates the bounding boxes of fruit objects using the text prompt “fruit”. These bounding boxes were provided as prompts to the segment anything model (SAM) to obtain segmentation masks, eliminating the need for manual annotation. To meet the real-time application requirements of embedded devices, a knowledge transfer method is proposed. The segmentation results of the large models were used as labels to train YOLO11n-seg, a considerably small model without
Zero-shot learning / Large vision model / Instance segmentation / Camellia oleifera fruit / Fruit location
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
|
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
|
| [42] |
|
| [43] |
|
| [44] |
|
| [45] |
|
| [46] |
|
/
| 〈 |
|
〉 |