MEGA Hub

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Authors

Do you know Shawn Li?You can claim authorship or link another user.Do you know Wei Yang?You can claim authorship or link another user.Do you know Jike Zhong?You can claim authorship or link another user.Do you know Jiate Li?You can claim authorship or link another user.Do you know Jiawei Yang?You can claim authorship or link another user.Do you know You Qin?You can claim authorship or link another user.Do you know Ryan Rossi?You can claim authorship or link another user.Do you know Franck Dernoncourt?You can claim authorship or link another user.Do you know Roger Zimmermann?You can claim authorship or link another user.Do you know Yue Wang?You can claim authorship or link another user.Do you know Zhengzhong Tu?You can claim authorship or link another user.Do you know Vicente Ordonez?You can claim authorship or link another user.Do you know Mohit Bansal?You can claim authorship or link another user.Do you know Yue Zhao?You can claim authorship or link another user.

Abstract

Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \ours{}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4times4 to 16times16), we find that zero-shot VLMs largely lack geometric reasoning: only one of five frontier models (GPT-5.5) exceeds random baseline on 4times4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves >97\% on 4times4, all models collapse on larger grids: GPT-5.5 drops from 70\% to near-random on 8times8, and even fine-tuned models fall below 5\% on 12times12. This ``scaling cliff'' suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. establishes scalable geometric reasoning as an open challenge for vision-language models.

Community

00