Segmented Summarization and Refinement:A Pipeline for Long-Document Analysis on Social Media

导出

摘要 Social media’s explosive growth has resulted in a massive influx of electronic documents influencing various facets of daily life.However,the enormous and complex nature of this content makes extracting valuable insights challenging.Long document summarization emerges as a pivotal technique in this context,serving to distill extensive texts into concise and comprehensible summaries.This paper presents a novel three-stage pipeline for effective long document summarization.The proposed approach combines unsupervised and supervised learning techniques,efficiently handling large document sets while requiring minimal computational resources.Our methodology introduces a unique process for forming semantic chunks through spectral dynamic segmentation,effectively reducing redundancy and repetitiveness in the summarization process.Contrary to previous methods,our approach aligns each semantic chunk with the entire summary paragraph,allowing the abstractive summarization model to process documents without truncation and enabling the summarization model to deduce missing information from other chunks.To enhance the summary generation,we utilize a sophisticated rewrite model based on Bidirectional and Auto-Regressive Transformers(BART),rearranging and reformulating summary constructs to improve their fluidity and coherence.Empirical studies conducted on the long documents from the Webis-TLDR-17 dataset demonstrate that our approach significantly enhances the efficiency of abstractive summarization transformers.The contributions of this paper thus offer significant advancements in the field of long document summarization,providing a novel and effective methodology for summarizing extensive texts in the context of social media.

作者 Guanghua Wang Priyanshi Garg Weili Wu

机构地区 the Department of Computer Science

出处《Journal of Social Computing》 EI 2024年第2期132-144,共13页 社会计算（英文）

关键词 long document summarization abstractive summarization text segmentation text alignment rewrite model spectral embedding

分类号 TP1 [自动化与计算机技术—控制理论与控制工程]

引文网络
相关文献

1葛云双,翟观文,李传海.数字音频技术在袋式除尘器维修中的应用研究[J].电声技术,2024,48(8):45-47.
2张岳琴,白丽霞,光明,朱镭,周浩,李秀辉,弓培慧,康娅楠.山西省百日咳发病的SARIMA模型预测[J].中国卫生统计,2024,41(4):551-554.
3Correction to“The exploration of cell population data in clinical use:Beyond infectious diseases”[J].iLABMED,2024,2(3):226-226.
4Yujie Ouyang,Min Zhang,Fangyang Zhan,Chunxia Li,Xianda Li,Fan Yan,Sen Xie,Qiwei Tong,Haoran Ge,Yong Liu,Rui Wang,Wei Liu,Xinfeng Tang.Intrinsically large effective mass and multi-valley band characteristics of n-type Bi_(2)eBi_(2)Te_(3)superlattice-like films[J].Journal of Materiomics,2024,10(3):716-724.
5张小龙,韩娇娇,王文波,李素兮,宁欣婷.基于ARIMA的延安市秋冬季最低温度预报研究[J].现代农业科技,2024(18):115-118.
6Xujie Gong,Ruichao Lei,Ruize Sun,Xue Jiang,Yanjing Su,Yu Yan.An ensemble learning strategy for multi-source hydrogen embrittlement data by introducing missing information[J].Materials Genome Engineering Advances,2024,2(2):145-157.
7Dandan Chu,Xingyue Yang,Jing Wang,Yan Zhou,Jin-Hua Gu,Jin Miao,Feng Wu,Fei Liu.Tau truncation in the pathogenesis of Alzheimer's disease:a narrative review[J].Neural Regeneration Research,2024,19(6):1221-1232. 被引量：3
8Vansh Sharma,Venkat Raman.A reliable knowledge processing framework for combustion science using foundation models[J].Energy and AI,2024,16(2):396-416.
9LIU Yong-shan.A Comparative Study of Artificial Intelligence and Translation Software in Chinese-English Translation:A Focus on Literary and Technical Texts[J].Journal of Literature and Art Studies,2024,14(9):815-820.
10陈佳俊,胡广漠,徐昕煜,薛波新.根治性前列腺切除术与近距离放疗治疗局限性前列腺癌的预后研究[J].现代泌尿生殖肿瘤杂志,2024,16(4):208-215.

Journal of Social Computing

2024年第2期

浏览历史

内容加载中请稍等...

Segmented Summarization and Refinement:A Pipeline for Long-Document Analysis on Social Media

相关作者

相关机构

相关主题

浏览历史