MEGA Hub

Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference

Authors

Do you know Yifan Lyu?You can claim authorship or link another user.Do you know Xinran Li?You can claim authorship or link another user.Do you know Jiaqi Qiao?You can claim authorship or link another user.Do you know Xiujuan Xu?You can claim authorship or link another user.

Abstract

Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and consequence across five fixed API models, six countries, and seven development-selected targets. In the primary post-review 72-record panel, opaque and disclosed-random labels each produced country-direction shifts of 0.214. Paired attenuation was 0.0003 (95% CI [-0.0157, 0.0166]). Verified country reduced Brier loss by 0.040 (95% CI [0.024, 0.056]), while random-label regret included zero. A non-overlapping mixed-coverage consistency panel retained positive disclosed-random movement and verified utility, while attenuation remained uncertain. On the selected targets, verified metadata was useful in both panels, but disclosure did not reliably attenuate random-label uptake. PROV-FORECAST contains 14,400 paired item-level probability distributions from the corrected panel.

Community

00

Publication notes

Author note
7 pages, 2 figures