AutoEval Done Right: Using Synthetic Data for Model Evaluation
Preprint 2024 en
Authors
PB
Pierre Boyeau
AA
Anastasios N. Angelopoulos
NY
Nir Yosef
Abstract
1 min read
The evaluation of machine learning models using human-labeled validation data can be expensive and time-consuming. AI-labeled synthetic data can be used to decrease the number of human annotations required for this purpose in a process called autoevaluation. We suggest efficient and statistically principled algorithms for this purpose that improve sample efficiency while remaining unbiased. These algorithms increase the effective human-labeled sample size by up to 50% on experiments with GPT-4.
Pietro Iaffaldano, Saverio D’Amico, Giuseppe Lucisano, Massimiliano Copetti, Tommaso Guerra, Maria A. Rocca, Francesco Patti, Giovanna De Luca, Diana Ferraro, Rocco Totaro, Vincenzo Brescia Morra, Giuseppe Salemi, Emilio Portaccio, Matteo Foschi, Matilde Inglese, Maria Gabriella Coniglio, Clara Grazia Chisari, Francesca Caputo, Damiano Paolicelli, Mario Alberto Battaglia, Matteo Della Porta, Victor Savevski,
Discussion(0)
No comments yet. Be the first to comment.