Text alignment is crucial to the accuracy of MT (Machine Translation) systems, some NLP (Natural Language Processing) tools or any other text processing tasks requiring bilingual data. This research proposes a lan...Text alignment is crucial to the accuracy of MT (Machine Translation) systems, some NLP (Natural Language Processing) tools or any other text processing tasks requiring bilingual data. This research proposes a language independent sentence alignment approach based on Polish (not position-sensitive language) to English experiments. This alignment approach was developed on the TED (Translanguage English Database) talks corpus, but can be used for any text domain or language pair. The proposed approach implements various heuristics for sentence recognition. Some of them value synonyms and semantic text structure analysis as a part of additional information. Minimization of data loss was ensured. The solution is compared to other sentence alignment implementations. Also an improvement in MT system score with text processed with the described tool is shown.展开更多
Social media’s explosive growth has resulted in a massive influx of electronic documents influencing various facets of daily life.However,the enormous and complex nature of this content makes extracting valuable insi...Social media’s explosive growth has resulted in a massive influx of electronic documents influencing various facets of daily life.However,the enormous and complex nature of this content makes extracting valuable insights challenging.Long document summarization emerges as a pivotal technique in this context,serving to distill extensive texts into concise and comprehensible summaries.This paper presents a novel three-stage pipeline for effective long document summarization.The proposed approach combines unsupervised and supervised learning techniques,efficiently handling large document sets while requiring minimal computational resources.Our methodology introduces a unique process for forming semantic chunks through spectral dynamic segmentation,effectively reducing redundancy and repetitiveness in the summarization process.Contrary to previous methods,our approach aligns each semantic chunk with the entire summary paragraph,allowing the abstractive summarization model to process documents without truncation and enabling the summarization model to deduce missing information from other chunks.To enhance the summary generation,we utilize a sophisticated rewrite model based on Bidirectional and Auto-Regressive Transformers(BART),rearranging and reformulating summary constructs to improve their fluidity and coherence.Empirical studies conducted on the long documents from the Webis-TLDR-17 dataset demonstrate that our approach significantly enhances the efficiency of abstractive summarization transformers.The contributions of this paper thus offer significant advancements in the field of long document summarization,providing a novel and effective methodology for summarizing extensive texts in the context of social media.展开更多
文摘Text alignment is crucial to the accuracy of MT (Machine Translation) systems, some NLP (Natural Language Processing) tools or any other text processing tasks requiring bilingual data. This research proposes a language independent sentence alignment approach based on Polish (not position-sensitive language) to English experiments. This alignment approach was developed on the TED (Translanguage English Database) talks corpus, but can be used for any text domain or language pair. The proposed approach implements various heuristics for sentence recognition. Some of them value synonyms and semantic text structure analysis as a part of additional information. Minimization of data loss was ensured. The solution is compared to other sentence alignment implementations. Also an improvement in MT system score with text processed with the described tool is shown.
文摘Social media’s explosive growth has resulted in a massive influx of electronic documents influencing various facets of daily life.However,the enormous and complex nature of this content makes extracting valuable insights challenging.Long document summarization emerges as a pivotal technique in this context,serving to distill extensive texts into concise and comprehensible summaries.This paper presents a novel three-stage pipeline for effective long document summarization.The proposed approach combines unsupervised and supervised learning techniques,efficiently handling large document sets while requiring minimal computational resources.Our methodology introduces a unique process for forming semantic chunks through spectral dynamic segmentation,effectively reducing redundancy and repetitiveness in the summarization process.Contrary to previous methods,our approach aligns each semantic chunk with the entire summary paragraph,allowing the abstractive summarization model to process documents without truncation and enabling the summarization model to deduce missing information from other chunks.To enhance the summary generation,we utilize a sophisticated rewrite model based on Bidirectional and Auto-Regressive Transformers(BART),rearranging and reformulating summary constructs to improve their fluidity and coherence.Empirical studies conducted on the long documents from the Webis-TLDR-17 dataset demonstrate that our approach significantly enhances the efficiency of abstractive summarization transformers.The contributions of this paper thus offer significant advancements in the field of long document summarization,providing a novel and effective methodology for summarizing extensive texts in the context of social media.