| 张志豪,孙志春,张硕,林歌歌,梁国强.语步结构引导的科技论文结构化自动摘要方法研究[J].数字图书馆论坛,2026,22(5):34~47 |
| 语步结构引导的科技论文结构化自动摘要方法研究 |
| Research on Move-Guided Structured Automatic Summarization for Scientific Papers |
| 投稿时间:2026-03-06 |
| DOI:10.3772/j.issn.1673–2286.2026.05.004 |
| 中文关键词: 科技论文;结构化摘要;语步结构;Seq2Seq框架;多任务学习 |
| 英文关键词: Scientific Paper; Structured Summaries; Move Structure; Seq2Seq Framework; Multi-task Learning |
| 基金项目:本研究得到国家自然科学基金委员会青年科学基金项目“大模型赋能的高价值专利个性化交易推荐方法与应用研究”(编号:72404020)、“颠覆性低碳技术识别及创新路径动态选择研究”(编号:72304023)及北京市自然科学基金面上项目“协同进化视角下研究前沿的涌现机制及识别方法研究”(编号:9232002)资助。 |
| 作者 | 单位 | | 张志豪 | 北京工业大学经济与管理学院 | | 孙志春 | 北京工业大学经济与管理学院 | | 张硕 | 北京工业大学经济与管理学院 | | 林歌歌 | 北京工业大学经济与管理学院 | | 梁国强 | 北京工业大学经济与管理学院 |
|
| 摘要点击次数: 30 |
| 全文下载次数: 35 |
| 中文摘要: |
| 科技论文篇幅较长,各章节研究要点分散。现有自动摘要方法仅依据句子权重、段落位置筛选文本,易丢失关键信息,导致摘要逻辑断裂;通用大模型生成摘要虽然行文流畅,但存在输出结构难以约束、事实描述失真等问题。为此,本研究提出语步引导的科技论文结构化摘要模型——Move-driven Summarization(MoverSum)。MoverSum采用Seq2Seq多任务学习框架,以细粒度语步为全局规划线索,融合文本语义、章节位置、语步标签构建多维特征,同步完成句子重要度评估与摘要语步序列预测,按照语步序列抽取关键语句,以生成条理清晰的结构化摘要。本研究基于arXiv和PubMed数据集开展综合实验,结果表明:MoverSum的各项ROUGE指标全面优于各类基线模型;人工评估进一步证实,该模型生成摘要具备更完整的内容覆盖度与更流畅的行文逻辑。本研究经消融实验、大模型对比实验证实,多维特征与语步规划是模型性能提升的核心,MoverSum对科技文本摘要场景适配度更高。本研究提出方法兼顾抽取式摘要的事实保真度与输出结构可控性,为科技论文结构化摘要任务提供全新语步感知建模思路。 |
| 英文摘要: |
| Scientific papers are typically lengthy, with key research points dispersed across various sections. Existing automatic summarization methods, which primarily rely on sentence weights and paragraph positions for text selection, often lose critical information, resulting in fragmented logic. While general-purpose large language models generate fluent summaries, they struggle to constrain output structures and frequently suffer from factual hallucinations. To address these challenges, this study proposes Move-driven Summarization (MoverSum), a move-guided structured summarization model for scientific papers. Built upon a multi-task learning Seq2Seq framework, MoverSum utilizes fine-grained rhetorical moves as a global planning guide. By integrating text semantics, section positions, and move labels into multi-dimensional features, the model simultaneously evaluates sentence salience and predicts the summary move sequence, extracting key sentences in accordance with the move sequence to generate well-organized, structured summaries. Comprehensive experiments conducted on the arXiv and PubMed datasets demonstrate that MoverSum consistently outperforms various baseline models across all ROUGE metrics. Furthermore, manual evaluation confirms that the generated summaries achieve superior content coverage and logical coherence. Ablation studies and comparative experiments with large language models verify that the multi-dimensional features and move-guided planning are the core drivers of performance
improvements, highlighting MoverSum’s high adaptability to scientific text summarization tasks. Ultimately, this proposed method balances the factual fidelity of extractive summarization with structural controllability, offering a novel move-aware modeling approach for the structured summarization of scientific papers. |
|
查看全文
查看/发表评论 下载PDF阅读器 |
| 关闭 |
|
|
|