MEGA Hub

Distinguishing case-mix from context heterogeneity in prognostic regression model synthesis settings

Authors

Do you know Max Behrens?You can claim authorship or link another user.Do you know Janis M. Nolde?You can claim authorship or link another user.Do you know Eleni Papakonstantinou?You can claim authorship or link another user.Do you know Gabriele Bellerino?You can claim authorship or link another user.Do you know Theodoros Evrenoglou?You can claim authorship or link another user.Do you know Angelika Rohde?You can claim authorship or link another user.Do you know Daiana Stolz?You can claim authorship or link another user.Do you know Moritz Hess?You can claim authorship or link another user.Do you know Harald Binder?You can claim authorship or link another user.

Abstract

Prognostic regression models often synthesize data from multiple sites, whether within a multi-site study, across federated settings, or in individual participant data meta-analysis. Here, a site is any data source, such as a hospital, registry, trial, or study, and need not be a physical center. Analysts must then decide whether one regression model represents all sites or whether site-specific models are needed. Established measures such as coefficient-level tau^2 quantify heterogeneity but do not distinguish its source. We focus on diagnosing whether coefficient heterogeneity reflects case-mix or site-specific context effects. Case-mix heterogeneity can arise when linear regression terms approximate multivariable non-linear relationships in populations with different covariate distributions. Contextual heterogeneity arises when comparable patients require different regression relationships across sites. We do this by fitting site-specific local regressions in a dimension-reduced space and partitioning the smoothed coefficient surfaces into a cross-site reference and site-specific deviations. An autoencoder and custom loss structure the latent space around local prognostic relationships. We then project this partition onto the outcome scale to derive observation- and site-level summaries. We demonstrate the approach on a COPD trial with two sites. In the three leading latent slope coordinates, coefficient-surface variation was predominantly contextual. The derived observation-level outcome-scale variance partition was case-mix-leading, whereas its between-site aggregation was concentrated in contextual differences rather than case-mix shifts. A permuted-site negative control assesses whether the contextual summary can arise when site labels carry no signal. This diagnostic distinction can inform whether joint or site-specific regression models should be evaluated.

Community

00

Publication notes

Author note
25 pages, 4 figures, 3 tables