如何用Python-Docx提取标题与表格间的多段文本?
Python-docx 处理Word文档:合并标题与表格间多段文本的解决方案
核心思路
问题根源在于之前的代码只遍历段落,没兼顾表格的位置顺序。解决的关键是按文档的实际元素顺序遍历(同时包含段落和表格),通过状态变量跟踪当前是否处于「收集标题与表格间文本」的阶段,从而把多段文本合并成一个块。
实现代码
from docx import Document from docx.enum.style import WD_STYLE_TYPE def process_word_doc(doc_path): doc = Document(doc_path) title_list = [] text_blocks = [] line_data = [] # 状态变量:记录当前标题、是否在收集文本、临时存储文本 current_title = None collecting_text = False current_text = "" # 遍历文档的所有顶层元素(段落+表格) for element in doc.element.body: # 判断是段落还是表格 if element.tag.endswith('p'): para = doc._body._element_to_object(element) # 判断是否是Heading 1且以'['开头 if para.style.style_type == WD_STYLE_TYPE.PARAGRAPH and para.style.name == 'Heading 1': if para.text.strip().startswith('['): # 处理上一个标题未匹配表格的剩余文本(按需保留或丢弃) if collecting_text and current_text: text_blocks.append(current_text.strip()) current_text = "" # 更新当前标题 current_title = para.text.strip() title_list.append(current_title) collecting_text = True else: # 普通段落,处于收集状态时合并文本 if collecting_text: current_text += para.text.strip() + "\n" elif element.tag.endswith('tbl'): table = doc._body._element_to_object(element) # 收集阶段结束,存入合并后的文本 if collecting_text: text_blocks.append(current_text.strip()) current_text = "" collecting_text = False # 提取表格内容存入line_data for row in table.rows: row_data = [cell.text.strip() for cell in row.cells] line_data.append(row_data) # 处理文档末尾未匹配表格的剩余文本 if collecting_text and current_text: text_blocks.append(current_text.strip()) return title_list, text_blocks, line_data # 示例调用 titles, texts, tables = process_word_doc("your_document.docx") # 后续可通过openpyxl等库将数据写入Excel
关键说明
- 遍历元素顺序:直接遍历
doc.element.body能保证段落和表格的顺序与文档完全一致,避免只遍历段落导致的表格位置丢失问题。 - 状态控制:
collecting_text标记确保只有在遇到目标标题后,才开始收集后续普通段落文本,直到遇到表格为止。 - 文本合并:用
current_text临时存储多段文本,遇到表格时一次性存入text_blocks,实现多段文本合并为一个块的需求。
内容的提问来源于stack exchange,提问作者Tom
相关产品推荐
相关产品推荐

