解决JSON数据处理中的NoneType无add属性错误及大数据量优化
问题解决:NoneType对象无add属性错误及大数据量优化方案
问题背景
给定两个输入数据结构:
first_df = {'data': {'historicalData': {'historical':[{'id': 52725940,'trades': [{'price': 99.94, 'size': 8}], 'product': None}, {'id': 52725941,'trades': None, 'product': {'price': 92.98,'executable': 18}}]}}} second_df = {'data': {'historicalData': {'historical':[{'id': 52725940,'trades': None, 'product': None}, {'id': 52725941,'trades': None, 'product': None}]}}}
需求:
- 当
product为None时,替换为{"price": None, "executable": None} - 当
trades为None时,替换为{'price': None, 'size': None}
用户编写的代码触发AttributeError: 'NoneType' object has no attribute 'add'错误,预期输出为:
id trades trades.price trades.size product product.price product.executable 0 52725940 {'price': None, 'size': None} None None {'price': None, 'executable': None} None None 1 52725941 {'price': None, 'size': None} None None {'price': None, 'executable': None} None None
错误原因
代码中使用i['product'].add(...)和i['trades'].add(...)的方式完全错误:当字段值为None时,None是Python的空类型对象,根本没有add方法。正确的做法是直接将None替换为目标字典,而非调用不存在的方法。
基础解决方法
修正后的代码直接遍历数据,替换指定字段的None值:
import pandas as pd second_df = {'data': {'historicalData': {'historical':[{'id': 52725940,'trades': None, 'product': None}, {'id': 52725941,'trades': None, 'product': None}]}}} historical = second_df['data']['historicalData']['historical'] # 遍历处理每个条目 for item in historical: # 替换product为默认字典(如果是None) if item['product'] is None: item['product'] = {"price": None, "executable": None} # 替换trades为默认字典(如果是None,注意原数据中trades可能是列表,需保留原有结构) if item['trades'] is None: item['trades'] = {'price': None, 'size': None} # 生成规范化的DataFrame output_df = pd.json_normalize(historical) print(output_df)
这段代码能同时处理first_df和second_df:遇到trades是列表的情况会保留原结构,仅当trades为None时替换为默认字典。
大数据量最优方案
当处理十万级以上的大数据集时,普通for循环效率偏低,推荐以下两种高效方案:
方案1:列表推导式(内存友好+速度快)
用自定义函数封装逻辑,结合列表推导式批量处理,底层基于C实现,比普通循环快数倍:
import pandas as pd def process_item(item): # 处理product字段 product = item['product'] if item['product'] is not None else {"price": None, "executable": None} # 处理trades字段:仅当为None时替换,保留原有列表结构 trades = item['trades'] if item['trades'] is not None else {'price': None, 'size': None} # 返回更新后的条目 return {**item, 'product': product, 'trades': trades} # 批量处理数据 historical = second_df['data']['historicalData']['historical'] processed_data = [process_item(item) for item in historical] # 生成DataFrame output_df = pd.json_normalize(processed_data)
如果数据集超大(百万级以上),可以用生成器表达式代替列表推导式,减少内存占用:
processed_data = (process_item(item) for item in historical) output_df = pd.json_normalize(processed_data)
方案2:Pandas管道式处理
如果数据已经加载到Pandas DataFrame中,用apply方法结合字段展开,适合超大规模数据集的流水线处理:
import pandas as pd # 原始数据转成DataFrame raw_df = pd.DataFrame(second_df['data']['historicalData']['historical']) # 定义字段处理函数 def fill_product_default(x): return {"price": None, "executable": None} if x is None else x def fill_trades_default(x): return {'price': None, 'size': None} if x is None else x # 批量处理字段 raw_df['product'] = raw_df['product'].apply(fill_product_default) raw_df['trades'] = raw_df['trades'].apply(fill_trades_default) # 展开嵌套字典生成最终结果 output_df = pd.json_normalize(raw_df.to_dict('records'))
内容的提问来源于stack exchange,提问作者NewUser
相关产品推荐
相关产品推荐

