使用python-docx提取Word不规则表格数据时遇错位缺失问题求助
Python-docx提取不规则Word表格的错位修复方案
问题背景
使用python-docx提取Word文档表格时,多数文件处理正常,但部分不规则表格出现单元格读取不全、数据错位问题。基础执行代码如下:
from docx.api import Document document = Document(file_path) [[cell.text for cell in row.cells] for row in document.tables[0].rows]
当前输出
执行上述代码后得到的错位输出:
[['', '', '', '', '', 'Nr.', '0b', '0c'], ['0d', '0e', '', '1b', '1c', '1d', '1e', ''], ['2b', '2c', '2d', '2e', '', '3b', '3c', '3d'], ['3e', '', '4b', '4c', '4d', '4e', '', '5b'], ['5c', '5d', '5e', '', '6b', '6c', '6d', '6e'], ['', '7b', '7c', '7d', '7e', '', '8b', '8c'], ['8d', '8e', 'foo', 'foo', 'foo', 'foo', 'foo', ''], ['', 'Click here to enter a date.', 'bar yes no\n\nfoobar\n', 'bar yes no\n\nfoobar\n', 'bar yes no\n\nfoobar\n', 'bar yes no\n\nfoobar\n', 'bar yes no\n\nfoobar\n'], [], [], [], []]
预期输出
期望得到的符合视觉结构的表格数据:
[['', '', '', '', ''], ['Nr.', '0b', '0c', '0d', '0e'], ['1.', '1b', '1c', '1d', '1e'], ['2.', '2b', '2c', '2d', '2e'], ['3.', '3b', '3c', '3d', '3e'], ['4.', '4b', '4c', '4d', '4e'], ['5.', '5b', '5c', '5d', '5e'], ['6.', '6b', '6c', '6d', '6e'], ['7.', '7b', '7c', '7d', '7e'], ['8.', '8b', '8c', '8d', '8e'], ['foo', 'foo', 'foo', 'foo', 'foo', ''], ['bar yes no\n\nfoobar\n', 'bar yes no\n\nfoobar\n', 'bar yes no\n\nfoobar\n', 'bar yes no\n\nfoobar\n', 'bar yes no\n\nfoobar\n']]
问题原因
这类问题并非普通跨列单元格导致,而是Word表格存在底层逻辑结构与视觉结构不一致的情况:
- 表格可能存在隐形的单元格拆分、跨行合并(非显性的跨列设置)
- Word的表格基于网格模型存储,视觉上的一行可能被拆分为多个逻辑行,或单元格存在隐藏的
v_merge(跨行)属性 - python-docx默认读取的是Word底层的逻辑单元格结构,而非视觉呈现的表格布局,因此出现数据错位
解决办法
通过解析单元格的grid_span(跨列)和v_merge(跨行)属性,手动重构符合视觉结构的表格:
from docx.api import Document def parse_irregular_table(table): # 计算表格实际的最大列数 max_cols = 0 row_spans = {} # 记录跨行单元格的剩余跨度:(行号, 列号): 剩余跨行数 # 第一遍遍历确定最大列数和跨行信息 for row_idx, row in enumerate(table.rows): col_idx = 0 for cell in row.cells: # 跳过被跨行单元格占据的列 while (row_idx, col_idx) in row_spans: row_spans[(row_idx, col_idx)] -= 1 if row_spans[(row_idx, col_idx)] == 0: del row_spans[(row_idx, col_idx)] col_idx += 1 colspan = cell._tc.grid_span # 获取跨行属性,无跨行则默认为1 rowspan = cell._tc.v_merge if cell._tc.v_merge is not None else 1 # 记录跨行信息 if rowspan > 1: for i in range(1, rowspan): row_spans[(row_idx + i, col_idx)] = rowspan - i # 更新最大列数 if col_idx + colspan > max_cols: max_cols = col_idx + colspan col_idx += colspan # 第二遍遍历构建视觉化的表格数据 result = [] row_spans = {} # 重置跨行记录 for row_idx, row in enumerate(table.rows): current_row = [''] * max_cols col_idx = 0 for cell in row.cells: # 跳过被跨行占据的列 while (row_idx, col_idx) in row_spans: row_spans[(row_idx, col_idx)] -= 1 if row_spans[(row_idx, col_idx)] == 0: del row_spans[(row_idx, col_idx)] col_idx += 1 colspan = cell._tc.grid_span rowspan = cell._tc.v_merge if cell._tc.v_merge is not None else 1 # 将单元格内容填充到对应的列位置 for c in range(colspan): if col_idx + c < max_cols: current_row[col_idx + c] = cell.text.strip() # 记录跨行信息 if rowspan > 1: for i in range(1, rowspan): row_spans[(row_idx + i, col_idx)] = rowspan - i col_idx += colspan # 过滤全空的行(按需保留) if any(cell.strip() for cell in current_row): result.append(current_row) return result # 使用示例 document = Document(file_path) fixed_table_data = parse_irregular_table(document.tables[0]) print(fixed_table_data)
代码说明
- 先遍历表格一次,计算表格实际的最大列数,并记录所有跨行单元格的位置和剩余跨度
- 再次遍历表格,根据跨行信息跳过被占据的单元格,将内容填充到对应的视觉位置
- 过滤掉全空的行,得到符合预期的表格结构
内容的提问来源于stack exchange,提问作者Ivar Eriksson
相关产品推荐
相关产品推荐

