如何在Pandas中将多个JSON字典合并为DataFrame?
合并JSON文件并转换为可筛选的DataFrame
错误原因
你遇到的ValueError: Expected object or value是因为参数使用错误:
- 你的JSON文件结构是以卡片ID为键的嵌套字典,并非
orient='records'要求的对象列表格式 lines=True用于读取每行一个JSON对象的JSONL格式文件,但你的文件是单个完整JSON对象,这个参数会导致解析失败
解决步骤
1. 正确读取单个JSON文件
针对你给出的JSON结构,需要用pd.DataFrame.from_dict()并指定orient='index',把卡片ID转为列,同时保留所有卡片信息:
import pandas as pd import json # 读取单个JSON文件并转为DataFrame def load_json_to_df(file_path): # 读取JSON为字典 with open(file_path, 'r', encoding='utf-8') as f: json_data = json.load(f) # 转换为DataFrame,orient='index'表示把字典的键作为行索引 df = pd.DataFrame.from_dict(json_data, orient='index') # 把行索引(卡片ID)转为单独的列,方便后续筛选 df = df.reset_index().rename(columns={'index': 'card_id'}) return df
2. 批量读取并合并所有JSON文件
遍历price-history目录下的所有JSON文件,读取后合并,同时添加source_file列标记数据来源:
from glob import glob # 获取所有JSON文件路径 arquivos = sorted(glob('price-history\\*.json')) # 批量读取并合并 dfs = [] for file in arquivos: df = load_json_to_df(file) # 添加来源文件列,区分不同文件的价格数据 df['source_file'] = file.split('\\')[-1] dfs.append(df) # 合并所有DataFrame todos_dados = pd.concat(dfs, ignore_index=True)
3. 验证与筛选
此时todos_dados就是可直接筛选的DataFrame,比如筛选价格大于0.01的记录:
# 筛选示例:价格大于0.01的卡片 filtered_df = todos_dados[todos_dados['price'] > 0.01] print(filtered_df)
4. 合并为单个JSON文件(可选)
如果需要把合并后的DataFrame转回JSON文件,可按原格式输出:
# 转为原格式(以card_id为键的嵌套字典) merged_json = todos_dados.set_index('card_id').to_dict(orient='index') # 保存为JSON文件 with open('merged_price_history.json', 'w', encoding='utf-8') as f: json.dump(merged_json, f, indent=2, ensure_ascii=False)
测试示例(用你给出的json1、json2模拟)
# 模拟两个JSON数据 json1 = { "105912": { "name": "Avatar - Tocasia, Dig Site Mentor", "cardset": "VAN", "rarity": "Rare", "foil": 0, "price": 0.05 }, "105911": { "name": "Avatar - Yotian Frontliner", "cardset": "VAN", "rarity": "Rare", "foil": 0, "price": 0.05 } } json2 = { "105912": { "name": "Avatar - Tocasia, Dig Site Mentor", "cardset": "VAN", "rarity": "Rare", "foil": 0, "price": 0.0007 }, "105911": { "name": "Avatar - Yotian Frontliner", "cardset": "VAN", "rarity": "Rare", "foil": 0, "price": 0.0007 } } # 模拟保存为文件 with open('price-history\\file1.json', 'w') as f: json.dump(json1, f) with open('price-history\\file2.json', 'w') as f: json.dump(json2, f) # 执行读取合并代码后,todos_dados的输出如下: # card_id name cardset rarity foil price source_file # 0 105912 Avatar - Tocasia, Dig Site Mentor VAN Rare 0 0.0500 file1.json # 1 105911 Avatar - Yotian Frontliner VAN Rare 0 0.0500 file1.json # 2 105912 Avatar - Tocasia, Dig Site Mentor VAN Rare 0 0.0007 file2.json # 3 105911 Avatar - Yotian Frontliner VAN Rare 0 0.0007 file2.json
内容的提问来源于stack exchange,提问作者Fabrício Gama
相关产品推荐
相关产品推荐

