Bracha Shapira

Senior Academic

Enhancing Data Annotation for Student Models

A Self-Training Approach with Large Language Models

Gilad Fuchs, Alex Nus, Yotam Eshel, Bracha Shapira

In this paper we introduce a methodology that utilizes Large Language Models (LLMs) for efficient data annotation which is essential for student models training. We specifically focus on classification tasks, assessing the impact of various annotation methods. We employ LLMs, such as GPT-4 and 7B parameter open-source models, as teacher models to facilitate the annotation process. We proceed to evaluate the efficiency of smaller, more deployment-efficient student models like BERT in production environments. We show two alternative methods. One using a cascade of LLMs distillation, starting from GPT-4 and following by 7B parameter LLM and BERT. And a second method, which shows the feasibility of annotating data without relying on proprietary model like GPT-4, by exclusively using a 7B parameter open-source model, leveraging self-training methods. This strategic use of smaller LLMs in conjunction with smaller student models presents an efficient and cost-effective solution for enhancing classification tasks, offering insights into the potential of utilizing large-scale language models in E-Commerce production settings.

Publication language English
Pages 2706-2712
Publication status Published - 23.05.2025

Keywords

Data Annotation
Large Language Models
Model Distillation

ASJC Scopus subject areas

Computer Networks and Communications
Software
Access to Document
10.1145/3701716.3717862
Other files and links
Link to publication in Scopus