You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python-Docx提取标题与表格间的多段文本?

Python-docx 处理Word文档:合并标题与表格间多段文本的解决方案

核心思路

问题根源在于之前的代码只遍历段落,没兼顾表格的位置顺序。解决的关键是按文档的实际元素顺序遍历(同时包含段落和表格),通过状态变量跟踪当前是否处于「收集标题与表格间文本」的阶段,从而把多段文本合并成一个块。

实现代码

from docx import Document
from docx.enum.style import WD_STYLE_TYPE

def process_word_doc(doc_path):
    doc = Document(doc_path)
    title_list = []
    text_blocks = []
    line_data = []
    
    # 状态变量:记录当前标题、是否在收集文本、临时存储文本
    current_title = None
    collecting_text = False
    current_text = ""
    
    # 遍历文档的所有顶层元素(段落+表格)
    for element in doc.element.body:
        # 判断是段落还是表格
        if element.tag.endswith('p'):
            para = doc._body._element_to_object(element)
            # 判断是否是Heading 1且以'['开头
            if para.style.style_type == WD_STYLE_TYPE.PARAGRAPH and para.style.name == 'Heading 1':
                if para.text.strip().startswith('['):
                    # 处理上一个标题未匹配表格的剩余文本(按需保留或丢弃)
                    if collecting_text and current_text:
                        text_blocks.append(current_text.strip())
                        current_text = ""
                    # 更新当前标题
                    current_title = para.text.strip()
                    title_list.append(current_title)
                    collecting_text = True
            else:
                # 普通段落,处于收集状态时合并文本
                if collecting_text:
                    current_text += para.text.strip() + "\n"
        elif element.tag.endswith('tbl'):
            table = doc._body._element_to_object(element)
            # 收集阶段结束,存入合并后的文本
            if collecting_text:
                text_blocks.append(current_text.strip())
                current_text = ""
                collecting_text = False
            # 提取表格内容存入line_data
            for row in table.rows:
                row_data = [cell.text.strip() for cell in row.cells]
                line_data.append(row_data)
    
    # 处理文档末尾未匹配表格的剩余文本
    if collecting_text and current_text:
        text_blocks.append(current_text.strip())
    
    return title_list, text_blocks, line_data

# 示例调用
titles, texts, tables = process_word_doc("your_document.docx")
# 后续可通过openpyxl等库将数据写入Excel

关键说明

  1. 遍历元素顺序:直接遍历doc.element.body能保证段落和表格的顺序与文档完全一致,避免只遍历段落导致的表格位置丢失问题。
  2. 状态控制:collecting_text标记确保只有在遇到目标标题后,才开始收集后续普通段落文本,直到遇到表格为止。
  3. 文本合并:用current_text临时存储多段文本,遇到表格时一次性存入text_blocks,实现多段文本合并为一个块的需求。

内容的提问来源于stack exchange,提问作者Tom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 05:12:35