Voxel51, the most powerful visual AI data platform, today released groundbreaking research showing that auto-labeling technology can achieve accuracy nearly equivalent to human labeling (up to 95%) while operating 5,000x faster than traditional annotation methods. Labeling costs can be reduced by up to 100,000x, potentially saving millions of dollars in AI development costs.
Read research on how zero-shot auto-labeling rivals human performance.
"Our research shows that data annotation no longer has to be a multi-million-dollar line item," said Jason Corso, Co-founder and Chief Science Officer at Voxel51. "While previous research has qualitatively claimed auto-labeling reduces annotation costs, our study provides concrete figures that have significant implications. The findings reflect the potential for a massive reduction in costs for data labeling, enabling AI developers to invest more of their budget and human workforce on more effective data curation, quality assurance, model and edge-case analysis, and strategic dataset expansion.”
Essential to powering computer vision, data labeling has traditionally been a tedious, costly, and slow process. To determine whether auto-labels alone could produce high-performing models in real-world scenarios, Voxel51’s Auto-Labeling Data for Object Detection study benchmarks leading foundation models, including YOLOE, YOLO-World, and Grounding DINO, across four widely-used datasets: Berkeley Deep Drive (BDD autonomous driving), Common Objects in Context (COCO), Large Vocabulary Instance Segmentation (LVIS high complexity), and Visual Object Classes (VOC general imagery). These datasets span basic object categories to challenging, long-tail distributions.
Using mean Average Precision (mAP), a key real-world metric for object detection accuracy, the study found that models trained solely on auto-labels performed just as well, and sometimes even better, than models trained on traditional human labels.
