MEGA Hub

Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

Authors

Do you know Han Wang?You can claim authorship or link another user.Do you know Alex Whitworth?You can claim authorship or link another user.Do you know Pak Ming Cheung?You can claim authorship or link another user.Do you know Zhenjie Zhang?You can claim authorship or link another user.Do you know Krishna Kamath?You can claim authorship or link another user.Do you know Xi Chen?You can claim authorship or link another user.Do you know Roberto Konow?You can claim authorship or link another user.Do you know Kurchi Subhra Hazra?You can claim authorship or link another user.

Abstract

Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging VLM-based labeling further unlocks opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effects (MDEs) in online experiment measurements.

Community

00

Publication notes

Author note
RecSys'26 Industry track
DOI
10.1145/3773078.3831891