含重复键的JSON转Pandas DataFrame时触发TypeError问题求助
解决JSON解码时的TypeError: expected string or buffer问题
你遇到的这个错误原因很清晰——JSONDecoder.decode()方法需要接收字符串或者字节缓冲区作为参数,但你直接把打开的文件对象f传进去了,它没法直接处理文件对象,所以抛出了类型错误。
下面给你两种简单的修正方案:
方法1:先读取文件内容为字符串
修改文件读取部分,先用f.read()把文件内容转换成字符串,再传给decode()方法:
from collections import OrderedDict from json import JSONDecoder import pandas as pd def make_unique(key, dct): counter = 0 unique_key = key while unique_key in dct: counter += 1 unique_key = '{}_{}'.format(key, counter) return unique_key def parse_object_pairs(pairs): dct = OrderedDict() for key, value in pairs: if key in dct: key = make_unique(key, dct) dct[key] = value return dct decoder = JSONDecoder(object_pairs_hook=parse_object_pairs) with open("file.json") as f: # 先读取文件内容为字符串再解码 json_content = f.read() obj = decoder.decode(json_content) # 将处理后的对象转换为DataFrame df = pd.DataFrame.from_dict(obj)
方法2:使用json.load()直接处理文件对象
其实json模块提供了load()方法,可以直接接收文件对象,并且同样支持传入object_pairs_hook参数,代码会更简洁:
from collections import OrderedDict import json import pandas as pd def make_unique(key, dct): counter = 0 unique_key = key while unique_key in dct: counter += 1 unique_key = '{}_{}'.format(key, counter) return unique_key def parse_object_pairs(pairs): dct = OrderedDict() for key, value in pairs: if key in dct: key = make_unique(key, dct) dct[key] = value return dct with open("file.json") as f: # 直接用json.load处理文件对象并应用自定义hook obj = json.load(f, object_pairs_hook=parse_object_pairs) # 转换为Pandas DataFrame df = pd.DataFrame.from_dict(obj)
两种方法都能解决你的问题,第二种更贴合json模块的常规用法。处理完obj之后,就可以顺利转换成DataFrame了。
内容的提问来源于stack exchange,提问作者Thabra
相关产品推荐
相关产品推荐

