You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas导出含Cell对象的DataFrame为JSON时属性丢失问题

自定义Cell类对象的Pandas DataFrame导出JSON不全问题

问题背景

我有一个包含自定义Cell类对象的Pandas DataFrame,Cell类包含cell_id、coordinates、data、empty等实例属性及相关方法。使用json.loads(df.to_json())导出为JSON时,仅能导出cell_id和data属性,其余属性未被包含。

相关代码

Cell类实现

class Cell:
    def __init__(self, cell_id, coordinates, data, empty):
        self.cell_id = cell_id
        self.coordinates = coordinates
        self.data = data
        self.empty = empty
    
    # 示例:若实现此方法会导致仅导出指定属性
    def __json__(self):
        return {"cell_id": self.cell_id, "data": self.data}

导出及复现代码

import pandas as pd
import json

# 创建测试DataFrame
sample_cells = [
    Cell(1, (15, 25), {"temp": 22}, False),
    Cell(2, (35, 45), {"temp": 28}, True)
]
df = pd.DataFrame({"cell": sample_cells})

# 执行导出
exported_json = json.loads(df.to_json(orient="records"))
print(exported_json)

导出结果

[{"cell": {"cell_id": 1, "data": {"temp": 22}}}, {"cell": {"cell_id": 2, "data": {"temp": 28}}}]

原因分析

Pandas的to_json()底层依赖Python标准库json模块完成序列化,默认逻辑如下:

  1. 若对象实现了__json__()方法,会直接调用该方法获取序列化内容;
  2. 若未实现,则尝试读取对象的__dict__属性(存储实例属性的字典)进行序列化。

你遇到的问题大概率是以下两种情况:

  • Cell类实现了__json__()方法,但仅返回了cell_id和data;
  • coordinates/empty不是直接的实例属性(比如是@property装饰的计算属性、或类属性),导致json模块无法从__dict__中读取到这些属性。

解决办法

方案1:修正__json__方法(如果已实现)

如果你的Cell类有__json__()方法,直接修改它包含所有需要导出的属性:

class Cell:
    def __init__(self, cell_id, coordinates, data, empty):
        self.cell_id = cell_id
        self.coordinates = coordinates
        self.data = data
        self.empty = empty
    
    def __json__(self):
        return {
            "cell_id": self.cell_id,
            "coordinates": self.coordinates,
            "data": self.data,
            "empty": self.empty
        }

方案2:自定义JSON编码器(适用于未实现__json__或属性为计算属性的情况)

编写自定义编码器,明确指定Cell对象的序列化规则:

import json
import pandas as pd

class CellEncoder(json.JSONEncoder):
    def default(self, obj):
        if isinstance(obj, Cell):
            # 手动指定需要序列化的所有属性,包括@property属性
            return {
                "cell_id": obj.cell_id,
                "coordinates": obj.coordinates,
                "data": obj.data,
                "empty": obj.empty
            }
        # 其他类型使用默认序列化逻辑
        return super().default(obj)

# 方式1:导出时指定编码器
json_str = df.to_json(orient="records", default_handler=lambda obj: json.dumps(obj, cls=CellEncoder))
result = json.loads(json_str)

# 方式2:先将DataFrame转为字典列表再序列化
df_dict = df.applymap(lambda x: CellEncoder().default(x) if isinstance(x, Cell) else x).to_dict(orient="records")
result = json.dumps(df_dict, indent=2)

方案3:直接提取实例属性(适用于属性都在__dict__中的情况)

如果coordinates和empty都是实例属性且存在于__dict__中,可以直接将Cell对象转为字典后再导出:

df_dict = df.applymap(lambda x: x.__dict__ if isinstance(x, Cell) else x).to_dict(orient="records")
json_str = json.dumps(df_dict, indent=2)

内容的提问来源于stack exchange,提问作者Yaniv Akiva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 22:36:34