如何用Python将带合并单元格的Word表格转为HTML?
解决Word表格转HTML的合并单元格问题
你当前的iter_unique_cells函数并未正确识别合并单元格,只是简单遍历所有单元格。要生成带合并单元格的HTML表格,核心是获取每个单元格的跨行列数(即rowspan和colspan属性),再对应生成HTML的<td>标签。
以下是修改后的完整代码,可正确处理合并单元格并输出HTML:
from docx import Document from docx.table import _Cell from docx.oxml.ns import qn def get_cell_span(cell: _Cell): # 获取单元格的跨行列数 tcPr = cell._tc.get_or_add_tcPr() # 处理colspan(跨列) gridSpan = tcPr.find(qn('w:gridSpan')) colspan = int(gridSpan.val) if gridSpan is not None else 1 # 处理rowspan(跨行) vMerge = tcPr.find(qn('w:vMerge')) if vMerge is not None: vMerge_val = vMerge.get(qn('w:val')) # "restart"表示是合并单元格的起始,统计跨行数 if vMerge_val == 'restart': rowspan = 1 row_idx = cell.parent._index table = cell.parent.parent for r in range(row_idx + 1, len(table.rows)): next_cell = table.rows[r].cells[cell._index] next_vMerge = next_cell._tc.get_or_add_tcPr().find(qn('w:vMerge')) if next_vMerge is not None and next_vMerge.get(qn('w:val')) != 'restart': rowspan += 1 else: break else: # 非起始单元格,无需输出内容 rowspan = 0 else: rowspan = 1 return colspan, rowspan def docx_table_to_html(doc_path): doc = Document(doc_path) html_tables = [] for table in doc.tables: html = ['<table border="1">'] # 记录需跳过的已合并单元格 skipped_cells = set() for row_idx, row in enumerate(table.rows): html.append(' <tr>') for cell_idx, cell in enumerate(row.cells): if (row_idx, cell_idx) in skipped_cells: continue colspan, rowspan = get_cell_span(cell) # 标记跨行后续单元格为跳过 if rowspan > 1: for r in range(row_idx + 1, row_idx + rowspan): skipped_cells.add((r, cell_idx)) # 提取单元格文本 cell_text = '\n'.join(p.text for p in cell.paragraphs).strip() # 生成<td>标签属性 td_attrs = [] if colspan > 1: td_attrs.append(f'colspan="{colspan}"') if rowspan > 1: td_attrs.append(f'rowspan="{rowspan}"') td_attr = ' '.join(td_attrs) if td_attr: html.append(f' <td {td_attr}>{cell_text}</td>') else: html.append(f' <td>{cell_text}</td>') html.append(' </tr>') html.append('</table>') html_tables.append('\n'.join(html)) return '\n\n'.join(html_tables) # 测试调用 html_output = docx_table_to_html("document.docx") print(html_output)
关键说明:
get_cell_span:读取Word底层XML属性,获取单元格的跨列数;跨行则需遍历后续行统计合并行数,因为Word仅在起始单元格标记restart,后续单元格仅标记延续。docx_table_to_html:遍历表格行与单元格,用skipped_cells集合跳过已合并的单元格,生成带colspan/rowspan属性的HTML表格。- 默认给HTML表格添加了
border="1",可按需调整样式。
内容的提问来源于stack exchange,提问作者user981
相关产品推荐
相关产品推荐

