%0 Conference Proceedings %T DSS: A Scalable and Efficient Stratified Sampling Algorithm for Large-Scale Datasets %+ National University of Defense Technology [China] %A Li, Minne %A Li, Dongsheng %A Shen, Siqi %A Zhang, Zhaoning %A Lu, Xicheng %Z Part 5: Data Processing and Big Data %< avec comité de lecture %( Lecture Notes in Computer Science %B 13th IFIP International Conference on Network and Parallel Computing (NPC) %C Xi'an, China %Y Guang R. Gao %Y Depei Qian %Y Xinbo Gao %Y Barbara Chapman %Y Wenguang Chen %I Springer International Publishing %3 Network and Parallel Computing %V LNCS-9966 %P 133-146 %8 2016-10-28 %D 2016 %R 10.1007/978-3-319-47099-3_11 %K Stratified sampling %K Aggregation %K Distributed processing %K Spark %Z Computer Science [cs]Conference papers %X Statistical analysis of aggregated records is widely used in various domains such as market research, sociological investigation and network analysis, etc. Stratified sampling (SS), which samples the population divided into distinct groups separately, is preferred in the practice for its high effectiveness and accuracy. In this paper, we propose a scalable and efficient algorithm named DSS, for SS to process large datasets. DSS executes all the sampling operations in parallel by calculating the exact subsample size for each partition according to the data distribution. We implement DSS on Spark, a big-data processing system, and we show through large-scale experiments that it can achieve lower data-transmission cost and higher efficiency than state-of-the-art methods with high sample representativeness. %G English %Z TC 10 %Z WG 10.3 %2 https://inria.hal.science/hal-01648006/document %2 https://inria.hal.science/hal-01648006/file/432484_1_En_11_Chapter.pdf %L hal-01648006 %U https://inria.hal.science/hal-01648006 %~ IFIP-LNCS %~ IFIP %~ IFIP-TC %~ IFIP-TC10 %~ IFIP-NPC %~ IFIP-WG10-3 %~ IFIP-LNCS-9966