You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决JSON数据处理中的NoneType无add属性错误及大数据量优化

问题解决:NoneType对象无add属性错误及大数据量优化方案

问题背景

给定两个输入数据结构:

first_df = {'data': {'historicalData': {'historical':[{'id': 52725940,'trades': [{'price': 99.94, 'size': 8}], 'product': None}, {'id': 52725941,'trades': None, 'product': {'price': 92.98,'executable': 18}}]}}}
second_df = {'data': {'historicalData': {'historical':[{'id': 52725940,'trades': None, 'product': None}, {'id': 52725941,'trades': None, 'product': None}]}}}

需求:

  • 当product为None时,替换为{"price": None, "executable": None}
  • 当trades为None时,替换为{'price': None, 'size': None}

用户编写的代码触发AttributeError: 'NoneType' object has no attribute 'add'错误,预期输出为:

id               trades  trades.price  trades.size                     product  product.price  product.executable
0  52725940  {'price': None, 'size': None}          None          None  {'price': None, 'executable': None}           None                None
1  52725941  {'price': None, 'size': None}          None          None  {'price': None, 'executable': None}           None                None

错误原因

代码中使用i['product'].add(...)和i['trades'].add(...)的方式完全错误:当字段值为None时,None是Python的空类型对象,根本没有add方法。正确的做法是直接将None替换为目标字典,而非调用不存在的方法。

基础解决方法

修正后的代码直接遍历数据,替换指定字段的None值:

import pandas as pd

second_df = {'data': {'historicalData': {'historical':[{'id': 52725940,'trades': None, 'product': None}, {'id': 52725941,'trades': None, 'product': None}]}}}
historical = second_df['data']['historicalData']['historical']

# 遍历处理每个条目
for item in historical:
    # 替换product为默认字典(如果是None)
    if item['product'] is None:
        item['product'] = {"price": None, "executable": None}
    # 替换trades为默认字典(如果是None,注意原数据中trades可能是列表,需保留原有结构)
    if item['trades'] is None:
        item['trades'] = {'price': None, 'size': None}

# 生成规范化的DataFrame
output_df = pd.json_normalize(historical)
print(output_df)

这段代码能同时处理first_df和second_df:遇到trades是列表的情况会保留原结构,仅当trades为None时替换为默认字典。

大数据量最优方案

当处理十万级以上的大数据集时,普通for循环效率偏低,推荐以下两种高效方案:

方案1:列表推导式(内存友好+速度快)

用自定义函数封装逻辑,结合列表推导式批量处理,底层基于C实现,比普通循环快数倍:

import pandas as pd

def process_item(item):
    # 处理product字段
    product = item['product'] if item['product'] is not None else {"price": None, "executable": None}
    # 处理trades字段:仅当为None时替换,保留原有列表结构
    trades = item['trades'] if item['trades'] is not None else {'price': None, 'size': None}
    # 返回更新后的条目
    return {**item, 'product': product, 'trades': trades}

# 批量处理数据
historical = second_df['data']['historicalData']['historical']
processed_data = [process_item(item) for item in historical]

# 生成DataFrame
output_df = pd.json_normalize(processed_data)

如果数据集超大(百万级以上),可以用生成器表达式代替列表推导式,减少内存占用:

processed_data = (process_item(item) for item in historical)
output_df = pd.json_normalize(processed_data)

方案2:Pandas管道式处理

如果数据已经加载到Pandas DataFrame中,用apply方法结合字段展开,适合超大规模数据集的流水线处理:

import pandas as pd

# 原始数据转成DataFrame
raw_df = pd.DataFrame(second_df['data']['historicalData']['historical'])

# 定义字段处理函数
def fill_product_default(x):
    return {"price": None, "executable": None} if x is None else x

def fill_trades_default(x):
    return {'price': None, 'size': None} if x is None else x

# 批量处理字段
raw_df['product'] = raw_df['product'].apply(fill_product_default)
raw_df['trades'] = raw_df['trades'].apply(fill_trades_default)

# 展开嵌套字典生成最终结果
output_df = pd.json_normalize(raw_df.to_dict('records'))

内容的提问来源于stack exchange,提问作者NewUser

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 08:44:54