MEGA Hub

A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

Authors

Do you know Katerina Papantoniou?You can claim authorship or link another user.Do you know Panagiotis Papadakos?You can claim authorship or link another user.Do you know Theodore Patkos?You can claim authorship or link another user.Do you know Dimitris Garefalakis?You can claim authorship or link another user.Do you know Nikos Vardakis?You can claim authorship or link another user.Do you know Dimitris Plexousakis?You can claim authorship or link another user.

Abstract

We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance judgments. We evaluate sparse (BM25), dense (sentence-transformers), hybrid, and LLM-assisted retrieval methods in this book-search setting. Multilingual embeddings outperform Greek-specific models, while hybrid retrieval performs best overall. A query-level analysis shows that BM25 excels at named-entity queries, while dense and hybrid methods improve natural-language, noisy, cross-lingual, and concept queries. Field-aware prompting has model-specific effects, while LLM TOC summarization improves TOC-only retrieval and LLM post-filtering improves early-stage retrieval at a high cost. Overall, CUP enables real-world evaluation of Greek retrieval across lexical, semantic, noisy, and cross-lingual queries.

Community

00

Publication notes

Author note
Preprint of a manuscript submitted to the 14th EETN Conference on Artificial Intelligence (SETN 2026)