Ultra-high-dimensional data with grouping structures arise naturally in many contemporary statistical problems,such as gene-wide association studies and the multi-factor analysis-of-variance(ANOVA).To address this iss...Ultra-high-dimensional data with grouping structures arise naturally in many contemporary statistical problems,such as gene-wide association studies and the multi-factor analysis-of-variance(ANOVA).To address this issue,we proposed a group screening method to do variables selection on groups of variables in linear models.This group screening method is based on a working independence,and sure screening property is also established for our approach.To enhance the finite sample performance,a data-driven thresholding and a two-stage iterative procedure are developed.To the best of our knowledge,screening for grouped variables rarely appeared in the literature,and this method can be regarded as an important and non-trivial extension of screening for individual variables.An extensive simulation study and a real data analysis demonstrate its finite sample performance.展开更多
基金supported by the National Natural Science Foundation of China(CN)(11571112)the National Social Science Foundation Key Program(17ZDA091)+1 种基金Natural Science Fund of Education Department of Anhui Province(KJ2013B233)the 111 Project of China(B14019).
文摘Ultra-high-dimensional data with grouping structures arise naturally in many contemporary statistical problems,such as gene-wide association studies and the multi-factor analysis-of-variance(ANOVA).To address this issue,we proposed a group screening method to do variables selection on groups of variables in linear models.This group screening method is based on a working independence,and sure screening property is also established for our approach.To enhance the finite sample performance,a data-driven thresholding and a two-stage iterative procedure are developed.To the best of our knowledge,screening for grouped variables rarely appeared in the literature,and this method can be regarded as an important and non-trivial extension of screening for individual variables.An extensive simulation study and a real data analysis demonstrate its finite sample performance.