如何修复合并JSON转CSV时出现的Pandas.errors.ParserError错误?
问题:合并多个JSON文件为CSV时触发ParserError错误
错误详情
pandas.errors.ParserError: Error tokenizing data. C error: Expected 1 fields in line 8, saw 2
现有代码
from glob import glob import pandas as pd import os.path path = '/mypath/*.json' files = glob(path) data = pd.concat((pd.read_json(file, error_bad_lines=False)) for file in files) data.to_csv(os.path.join('/path where I want to put the csv file','yourfilename.csv'))
JSON文件结构
{ "data":[ { "attachments":"3_1630878910203273218", "id":"1630878914296913926" }, { "attachments":"4839577767895768", "id":"18945865497657868" } ], "includes":{ "media":[ { "media_key":"3_1630878910203273218", "type":"photo", "url":"https://pbs.twimg.com/media/FqILumJXoAIOzy4.jpg" } ] }, "meta":{ "newest_id":"1630878914296913926", "oldest_id":"1627974617309519873", "result_count":38 } }
期望输出CSV结构
| attachments | id |
|---|---|
| 3_1630878910203273218 | 1630878914296913926 |
| 4839577767895768 | 18945865497657868 |
解决方案
错误原因
- 原代码直接读取整个JSON文件,会将顶层的
data、includes、meta作为DataFrame列,导致嵌套数据结构混乱,触发解析错误。 pd.concat的语法存在隐患,生成器表达式外层缺少括号,易引发结构匹配问题。- 未指定仅提取JSON中的
data数组部分,不符合需求。
修正后的代码
from glob import glob import pandas as pd import os.path # 定义路径 json_path = '/mypath/*.json' output_path = os.path.join('/path where I want to put the csv file', 'yourfilename.csv') # 收集所有data部分的DataFrame dfs = [] for file in glob(json_path): # 读取JSON并提取data字段 json_data = pd.read_json(file) df = pd.DataFrame(json_data['data'].tolist()) dfs.append(df) # 合并所有DataFrame combined_data = pd.concat(dfs, ignore_index=True) # 保存为CSV combined_data.to_csv(output_path, index=False)
关键说明
- 通过
json_data['data'].tolist()将嵌套的data数组转换为可直接生成DataFrame的列表,精准提取目标数据。 - 使用
ignore_index=True重置合并后的索引,避免重复索引问题。 - 添加
index=False避免将DataFrame索引写入CSV文件,完全匹配期望输出格式。
内容的提问来源于stack exchange,提问作者LZL
相关产品推荐
相关产品推荐

