如何用Python docx读取并处理Word中无法识别的表格数据
问题描述
我有一个docx文件,其中的表格无法被doc.tables识别,相关问题在python-docx的开源仓库中也存在未解决的记录。以下是我的测试代码,需要解决方案:
from docx import Document import pandas as pd doc = Document("non_readable_table.docx") print(doc.tables) def iter_tables(block_item_container): """Recursively generate all tables in `block_item_container`.""" for t in block_item_container.tables: yield t for row in t.rows: for cell in row.cells: yield from iter_tables(cell) dfs = [] for t in iter_tables(doc): table = t df = [['' for i in range(len(table.columns))] for j in range(len(table.rows))] for i, row in enumerate(table.rows): for j, cell in enumerate(row.cells): if cell.text: df[i][j] = cell.text.replace('\n', '') dfs.append(pd.DataFrame(df)) print(dfs)
解决方案
这种情况通常是因为表格并非Word原生表格,而是用文本框、线条模拟的伪表格,python-docx仅能识别标准<w:tbl>标签定义的表格,对模拟表格无法直接识别。可尝试以下方案:
解析docx的XML结构:
docx是zip压缩包,内部word/document.xml存储核心内容。直接读取该XML,根据文档实际结构定位表格相关元素(比如文本排版规律、线条元素)提取数据,示例框架:from zipfile import ZipFile import xml.etree.ElementTree as ET with ZipFile("non_readable_table.docx", 'r') as zf: with zf.open('word/document.xml') as f: tree = ET.parse(f) root = tree.getroot() # 根据XML标签特征编写提取逻辑,需匹配实际文档结构用OCR工具提取结构化数据:
将docx转为图片后,使用OCR工具(如pytesseract)识别内容,再通过表格识别库(如tabula-py)提取结构化表格数据。更换处理库提取纯文本重构表格:
使用docx2txt或textract提取纯文本,根据文本的缩进、换行、分隔符规律重构表格:import docx2txt text = docx2txt.process("non_readable_table.docx") rows = [row.strip() for row in text.split('\n') if row.strip()] table_data = [row.split('\t') for row in rows]手动转换为原生表格:
若文档数量少,可手动打开Word文件,选中模拟表格区域,右键选择「转换为表格」,将伪表格转为原生结构后,即可用python-docx正常识别。
内容的提问来源于stack exchange,提问作者xie186
相关产品推荐
相关产品推荐

