You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复Python提取DOCX表格生成Pandas DataFrame的格式问题?

解决方案:修复Word表格转DataFrame的格式问题并对比差异

你的方案完全可行——将Word表格转JSON后用Pandas对比是处理原位更新表格差异的高效方式。针对simplify_docx提取的JSON结构导致DataFrame格式异常的问题,只需先扁平化嵌套的表格数据,再生成标准DataFrame即可,以下是具体步骤:

1. 解析simplify_docx输出,转换为标准二维列表

simplify_docx提取的表格数据是嵌套字典结构(tables -> rows -> cells -> text),直接转DataFrame会把整个行/列嵌套结构当成单个字段。我们需要先提取每个单元格的纯文本,转换成行×列的二维列表:

import simplify_docx
import pandas as pd

def docx_to_dataframe(doc_path):
    # 解析Word文档
    doc_content = simplify_docx.parse_docx(doc_path)
    # 提取目标表格(假设是文档中第一个表格,多表格可调整索引或匹配特征)
    target_table = doc_content['tables'][0]
    # 扁平化表格数据:遍历每行,提取单元格文本
    flattened_data = []
    for row in target_table['rows']:
        row_cells = [cell['text'] for cell in row['cells']]
        flattened_data.append(row_cells)
    # 用第一行作为表头,剩余行作为数据生成DataFrame
    df = pd.DataFrame(flattened_data[1:], columns=flattened_data[0])
    return df

2. 生成两个文档的DataFrame并对比差异

调用上述函数生成两个文档的DataFrame后,即可用Pandas内置工具对比差异:

# 生成两个文档的DataFrame
df_file1 = docx_to_dataframe('file-1.docx')
df_file2 = docx_to_dataframe('file-2.docx')

# 方法1:用Pandas内置的compare生成结构化差异报告
diff_report = df_file1.compare(df_file2)
print("结构化差异报告:")
print(diff_report)

# 方法2:单元格级逐行逐列对比(适合需要精准定位差异位置的场景)
print("\n单元格级差异详情:")
for col in df_file1.columns:
    for idx in df_file1.index:
        val1 = df_file1.loc[idx, col]
        val2 = df_file2.loc[idx, col]
        if val1 != val2:
            print(f"位置({idx+1}, {col}):原内容='{val1}' | 修改后='{val2}'")

注意事项

  • 多表格处理:如果文档包含多个表格,可通过表格的行数、表头内容等特征匹配目标表格,避免取错索引。
  • 带格式的单元格:如果单元格包含加粗、换行等格式,simplify_docx会在cell['runs']中存储格式信息,若只需纯文本则继续使用cell['text']即可。
  • 行/列增删处理:若原位更新包含行/列的添加或删除,需先对两个DataFrame做索引对齐(比如用pd.merge或reindex),再进行对比。

内容的提问来源于stack exchange,提问作者Vaibhav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 00:32:42