You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas去除ODS表格中的单元格批注(Insert Comment类型)

如何读取ODS表格时忽略单元格批注,只获取原始内容

问题场景

你有一个带多工作表的ODS文件,表头行的单元格添加了右键插入的批注(比如第一个单元格的批注是“Inserted comment”),用Pandas的read_excel配合odf引擎读取时,批注内容会和单元格文本拼接、拆分混乱,导致表头变成类似['commentfield_name', 'alt_names', 'type']的错误结果,而你只需要原始表头['field_name', 'alt_names', 'type']。

解决方案:直接用odfpy解析ODS结构,跳过批注节点

Pandas的封装读取会错误解析批注元数据,所以我们直接操作ODS的XML结构,只提取单元格的纯文本内容,完全避开批注节点。

步骤1:安装依赖

pip install odfpy

步骤2:编写解析代码

from odf.opendocument import load
from odf.table import Table, TableRow, TableCell

def read_ods_without_comments(file_path, sheet_name):
    # 加载ODS文件
    doc = load(file_path)
    # 遍历找到目标工作表
    for sheet in doc.getElementsByType(Table):
        if sheet.getAttribute('name') == sheet_name:
            rows_content = []
            # 遍历每一行
            for row in sheet.getElementsByType(TableRow):
                current_row = []
                # 遍历行内每个单元格
                for cell in row.getElementsByType(TableCell):
                    # 只提取单元格的纯文本节点内容,跳过批注相关节点
                    cell_text = ''.join([node.data for node in cell.childNodes if node.nodeType == node.TEXT_NODE])
                    # 处理单元格合并的情况(重复填充内容)
                    repeat_count = cell.getAttribute('numbercolumnsrepeated')
                    if repeat_count:
                        current_row.extend([cell_text] * int(repeat_count))
                    else:
                        current_row.append(cell_text)
                rows_content.append(current_row)
            return rows_content
    # 没找到目标工作表时抛出异常
    raise ValueError(f"找不到名为 {sheet_name} 的工作表")

# 调用示例
full_data = read_ods_without_comments('file.ods', 'x')
# 获取表头行
header_row = full_data[0]
print(header_row)  # 输出:['field_name', 'alt_names', 'type']

可选:转成Pandas DataFrame

如果需要后续用Pandas处理数据,可以把解析结果转成DataFrame:

import pandas as pd

ods_data = read_ods_without_comments('file.ods', 'x')
# 第一行作为表头,剩余行作为数据
df = pd.DataFrame(ods_data[1:], columns=ods_data[0])

为什么这个方法有效?

ODS文件里的批注是独立的<annotation>节点,和单元格的文本节点<text:p>是分离的。Pandas的read_excel用odf引擎时,底层错误地把批注的创建日期、内容等元数据拼接到了单元格文本中;而我们的代码只提取单元格内的纯文本节点,完全不处理批注相关的节点,所以不会出现内容混乱的问题。

内容的提问来源于stack exchange,提问作者NetAlien

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 07:27:25