You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用python-docx提取Word不规则表格数据时遇错位缺失问题求助

Python-docx提取不规则Word表格的错位修复方案

问题背景

使用python-docx提取Word文档表格时,多数文件处理正常,但部分不规则表格出现单元格读取不全、数据错位问题。基础执行代码如下:

from docx.api import Document
document = Document(file_path)    
[[cell.text for cell in row.cells] for row in document.tables[0].rows]

当前输出

执行上述代码后得到的错位输出:

[['', '', '', '', '', 'Nr.', '0b', '0c'],
 ['0d', '0e', '', '1b', '1c', '1d', '1e', ''],
 ['2b', '2c', '2d', '2e', '', '3b', '3c', '3d'],
 ['3e', '', '4b', '4c', '4d', '4e', '', '5b'],
 ['5c', '5d', '5e', '', '6b', '6c', '6d', '6e'],
 ['', '7b', '7c', '7d', '7e', '', '8b', '8c'],
 ['8d', '8e', 'foo', 'foo', 'foo', 'foo', 'foo', ''],
 ['',
  'Click here to enter a date.',
  'bar    yes    no\n\nfoobar\n',
  'bar    yes    no\n\nfoobar\n',
  'bar    yes    no\n\nfoobar\n',
  'bar    yes    no\n\nfoobar\n',
  'bar    yes    no\n\nfoobar\n'],
 [],
 [],
 [],
 []]

预期输出

期望得到的符合视觉结构的表格数据:

[['', '', '', '', ''], 
 ['Nr.', '0b', '0c', '0d', '0e'],
 ['1.', '1b', '1c', '1d', '1e'],
 ['2.', '2b', '2c', '2d', '2e'],
 ['3.', '3b', '3c', '3d', '3e'],
 ['4.', '4b', '4c', '4d', '4e'],
 ['5.', '5b', '5c', '5d', '5e'],
 ['6.', '6b', '6c', '6d', '6e'],
 ['7.', '7b', '7c', '7d', '7e'], 
 ['8.', '8b', '8c', '8d', '8e'],
 ['foo', 'foo', 'foo', 'foo', 'foo', ''],
 ['bar    yes    no\n\nfoobar\n',
  'bar    yes    no\n\nfoobar\n',
  'bar    yes    no\n\nfoobar\n',
  'bar    yes    no\n\nfoobar\n',
  'bar    yes    no\n\nfoobar\n']]

问题原因

这类问题并非普通跨列单元格导致,而是Word表格存在底层逻辑结构与视觉结构不一致的情况:

  • 表格可能存在隐形的单元格拆分、跨行合并(非显性的跨列设置)
  • Word的表格基于网格模型存储,视觉上的一行可能被拆分为多个逻辑行,或单元格存在隐藏的v_merge(跨行)属性
  • python-docx默认读取的是Word底层的逻辑单元格结构,而非视觉呈现的表格布局,因此出现数据错位

解决办法

通过解析单元格的grid_span(跨列)和v_merge(跨行)属性,手动重构符合视觉结构的表格:

from docx.api import Document

def parse_irregular_table(table):
    # 计算表格实际的最大列数
    max_cols = 0
    row_spans = {}  # 记录跨行单元格的剩余跨度:(行号, 列号): 剩余跨行数

    # 第一遍遍历确定最大列数和跨行信息
    for row_idx, row in enumerate(table.rows):
        col_idx = 0
        for cell in row.cells:
            # 跳过被跨行单元格占据的列
            while (row_idx, col_idx) in row_spans:
                row_spans[(row_idx, col_idx)] -= 1
                if row_spans[(row_idx, col_idx)] == 0:
                    del row_spans[(row_idx, col_idx)]
                col_idx += 1
            
            colspan = cell._tc.grid_span
            # 获取跨行属性,无跨行则默认为1
            rowspan = cell._tc.v_merge if cell._tc.v_merge is not None else 1
            
            # 记录跨行信息
            if rowspan > 1:
                for i in range(1, rowspan):
                    row_spans[(row_idx + i, col_idx)] = rowspan - i
            
            # 更新最大列数
            if col_idx + colspan > max_cols:
                max_cols = col_idx + colspan
            col_idx += colspan

    # 第二遍遍历构建视觉化的表格数据
    result = []
    row_spans = {}  # 重置跨行记录
    for row_idx, row in enumerate(table.rows):
        current_row = [''] * max_cols
        col_idx = 0
        for cell in row.cells:
            # 跳过被跨行占据的列
            while (row_idx, col_idx) in row_spans:
                row_spans[(row_idx, col_idx)] -= 1
                if row_spans[(row_idx, col_idx)] == 0:
                    del row_spans[(row_idx, col_idx)]
                col_idx += 1
            
            colspan = cell._tc.grid_span
            rowspan = cell._tc.v_merge if cell._tc.v_merge is not None else 1
            
            # 将单元格内容填充到对应的列位置
            for c in range(colspan):
                if col_idx + c < max_cols:
                    current_row[col_idx + c] = cell.text.strip()
            
            # 记录跨行信息
            if rowspan > 1:
                for i in range(1, rowspan):
                    row_spans[(row_idx + i, col_idx)] = rowspan - i
            col_idx += colspan
        
        # 过滤全空的行(按需保留)
        if any(cell.strip() for cell in current_row):
            result.append(current_row)
    return result

# 使用示例
document = Document(file_path)
fixed_table_data = parse_irregular_table(document.tables[0])
print(fixed_table_data)

代码说明

  1. 先遍历表格一次,计算表格实际的最大列数,并记录所有跨行单元格的位置和剩余跨度
  2. 再次遍历表格,根据跨行信息跳过被占据的单元格,将内容填充到对应的视觉位置
  3. 过滤掉全空的行,得到符合预期的表格结构

内容的提问来源于stack exchange,提问作者Ivar Eriksson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 01:57:04