TY - JOUR
T1 - Novel Distance Regression for Repeated Outcomes With Missing Data
T2 - Applications to Longitudinal and Crossover Studies of Microbiome Beta-Diversity
AU - Liu, Jinyuan
AU - Xu, Ke
AU - Ferguson, Jane F.
AU - Kang, Kaidi
AU - Wang, Yue
AU - Qiu, Yuqi
AU - Shao, Lucy
AU - Tu, Shengjia
AU - Nguyen, Tanya T.
AU - Lin, Tuo
AU - Zhang, Xinlian
N1 - Publisher Copyright:
© 2026 The Author(s). Statistics in Medicine published by John Wiley & Sons Ltd.
PY - 2026/7
Y1 - 2026/7
N2 - The human microbiome plays a crucial role in health, but understanding its dynamic relationship with the host requires regular monitoring. Beyond challenges such as high dimensionality and sparsity, additional complexities arise, particularly within-cluster correlation from repeated measures and pervasive missing data. To address these issues, we develop Edger, a novel distance regression method for modeling community-level beta-diversity dynamics and their interactions with treatment or host physiology. By focusing on beta-diversity, a distance metric between microbial profiles, Edger (Ensembled semiparametric distance-based generalized estimation for repeated outcomes) directly models these distances as repeated outcomes, yielding interpretable coefficients and enabling a covariate batching strategy to mitigate omitted variable bias. Our semiparametric inference framework eliminates the need for time-consuming permutation tests, distinguishes between-cluster heterogeneity from within-cluster fluctuations, and allows flexible specification of working correlation structures. To handle missing data, we assume a missing-at-random (MAR) mechanism and incorporate a between-subject propensity score in the repeated distance regression to provide seamless joint inference, ensuring robust variance estimation without casewise deletion. Additionally, we introduce an algorithm to generate synthetic data from real-world microbial counts while preserving their zero-inflated and correlated nature. Edger demonstrates superior inferential power and computational efficiency through our numerical studies and real-world applications, making it a valuable tool for uncovering microbiome-host interactions and advancing multi-omics data integration.
AB - The human microbiome plays a crucial role in health, but understanding its dynamic relationship with the host requires regular monitoring. Beyond challenges such as high dimensionality and sparsity, additional complexities arise, particularly within-cluster correlation from repeated measures and pervasive missing data. To address these issues, we develop Edger, a novel distance regression method for modeling community-level beta-diversity dynamics and their interactions with treatment or host physiology. By focusing on beta-diversity, a distance metric between microbial profiles, Edger (Ensembled semiparametric distance-based generalized estimation for repeated outcomes) directly models these distances as repeated outcomes, yielding interpretable coefficients and enabling a covariate batching strategy to mitigate omitted variable bias. Our semiparametric inference framework eliminates the need for time-consuming permutation tests, distinguishes between-cluster heterogeneity from within-cluster fluctuations, and allows flexible specification of working correlation structures. To handle missing data, we assume a missing-at-random (MAR) mechanism and incorporate a between-subject propensity score in the repeated distance regression to provide seamless joint inference, ensuring robust variance estimation without casewise deletion. Additionally, we introduce an algorithm to generate synthetic data from real-world microbial counts while preserving their zero-inflated and correlated nature. Edger demonstrates superior inferential power and computational efficiency through our numerical studies and real-world applications, making it a valuable tool for uncovering microbiome-host interactions and advancing multi-omics data integration.
KW - U-statistics
KW - between-subject outcome
KW - feature aggregation
KW - missing at random
KW - semiparametric inference
KW - weighted estimating equation
UR - https://www.scopus.com/pages/publications/105043837359
U2 - 10.1002/sim.70654
DO - 10.1002/sim.70654
M3 - 文章
AN - SCOPUS:105043837359
SN - 0277-6715
VL - 45
JO - Statistics in Medicine
JF - Statistics in Medicine
IS - 15-17
M1 - e70654
ER -