You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将带合并单元格的Word表格转为HTML?

解决Word表格转HTML的合并单元格问题

你当前的iter_unique_cells函数并未正确识别合并单元格,只是简单遍历所有单元格。要生成带合并单元格的HTML表格,核心是获取每个单元格的跨行列数(即rowspan和colspan属性),再对应生成HTML的<td>标签。

以下是修改后的完整代码,可正确处理合并单元格并输出HTML:

from docx import Document
from docx.table import _Cell
from docx.oxml.ns import qn

def get_cell_span(cell: _Cell):
    # 获取单元格的跨行列数
    tcPr = cell._tc.get_or_add_tcPr()
    # 处理colspan(跨列)
    gridSpan = tcPr.find(qn('w:gridSpan'))
    colspan = int(gridSpan.val) if gridSpan is not None else 1
    # 处理rowspan(跨行)
    vMerge = tcPr.find(qn('w:vMerge'))
    if vMerge is not None:
        vMerge_val = vMerge.get(qn('w:val'))
        # "restart"表示是合并单元格的起始,统计跨行数
        if vMerge_val == 'restart':
            rowspan = 1
            row_idx = cell.parent._index
            table = cell.parent.parent
            for r in range(row_idx + 1, len(table.rows)):
                next_cell = table.rows[r].cells[cell._index]
                next_vMerge = next_cell._tc.get_or_add_tcPr().find(qn('w:vMerge'))
                if next_vMerge is not None and next_vMerge.get(qn('w:val')) != 'restart':
                    rowspan += 1
                else:
                    break
        else:
            # 非起始单元格,无需输出内容
            rowspan = 0
    else:
        rowspan = 1
    return colspan, rowspan

def docx_table_to_html(doc_path):
    doc = Document(doc_path)
    html_tables = []
    for table in doc.tables:
        html = ['<table border="1">']
        # 记录需跳过的已合并单元格
        skipped_cells = set()
        for row_idx, row in enumerate(table.rows):
            html.append('  <tr>')
            for cell_idx, cell in enumerate(row.cells):
                if (row_idx, cell_idx) in skipped_cells:
                    continue
                colspan, rowspan = get_cell_span(cell)
                # 标记跨行后续单元格为跳过
                if rowspan > 1:
                    for r in range(row_idx + 1, row_idx + rowspan):
                        skipped_cells.add((r, cell_idx))
                # 提取单元格文本
                cell_text = '\n'.join(p.text for p in cell.paragraphs).strip()
                # 生成<td>标签属性
                td_attrs = []
                if colspan > 1:
                    td_attrs.append(f'colspan="{colspan}"')
                if rowspan > 1:
                    td_attrs.append(f'rowspan="{rowspan}"')
                td_attr = ' '.join(td_attrs)
                if td_attr:
                    html.append(f'    <td {td_attr}>{cell_text}</td>')
                else:
                    html.append(f'    <td>{cell_text}</td>')
            html.append('  </tr>')
        html.append('</table>')
        html_tables.append('\n'.join(html))
    return '\n\n'.join(html_tables)

# 测试调用
html_output = docx_table_to_html("document.docx")
print(html_output)

关键说明:

  • get_cell_span:读取Word底层XML属性,获取单元格的跨列数;跨行则需遍历后续行统计合并行数,因为Word仅在起始单元格标记restart,后续单元格仅标记延续。
  • docx_table_to_html:遍历表格行与单元格,用skipped_cells集合跳过已合并的单元格,生成带colspan/rowspan属性的HTML表格。
  • 默认给HTML表格添加了border="1",可按需调整样式。

内容的提问来源于stack exchange,提问作者user981

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 00:11:01