使用Pandas分块读取JSON文件并生成合并CSV时遭遇ValueError错误求助
解决
ValueError: Expected object or value及代码语法问题 咱们一步步拆解你遇到的问题,这个错误主要来自两个核心原因:JSON文件存在无效行,加上代码里的布尔筛选语法写错了,一起来修复:
1. 修复JSON读取时的无效行问题
当用pd.read_json(lines=True)读取大JSON文件时,如果某一行不是合法的JSON格式(比如引号不闭合、括号缺失),就会触发ValueError: Expected object or value。咱们可以自定义一个带错误处理的迭代器,跳过这些坏行:
import pandas as pd import json def read_json_safely(file_path, chunksize=1000000): with open(file_path, 'r', encoding='utf-8') as f: chunk = [] for line_num, line in enumerate(f, 1): try: chunk.append(json.loads(line)) if len(chunk) >= chunksize: yield pd.DataFrame(chunk) chunk = [] except json.JSONDecodeError: print(f"跳过无效行 {line_num}: {line[:50]}...") if chunk: yield pd.DataFrame(chunk) # 用自定义迭代器替代原pd.read_json的chunk读取 df_reader = read_json_safely('Clothing_Shoes_and_Jewelry.json', chunksize=1000000)
2. 修正布尔筛选的语法错误
你的代码里所有的评分筛选都写错了!比如new_df[new_df['overall' == 1]]应该是new_df[new_df['overall'] == 1]——你把比较符号放错了位置,导致程序试图用False('overall' ==1的结果)当列名,直接会触发KeyError。修正后的筛选代码:
counter = 1 for chunk in df_reader: new_df = chunk[['overall', 'reviewText','summary']].copy() # 直接切片更简洁 # 修正后的布尔筛选逻辑 new_df1 = new_df[new_df['overall'] == 1].sample(4000, replace=False) new_df2 = new_df[new_df['overall'] == 2].sample(4000, replace=False) new_df3 = new_df[new_df['overall'] == 4].sample(4000, replace=False) new_df4 = new_df[new_df['overall'] == 5].sample(4000, replace=False) new_df5 = new_df[new_df['overall'] == 3].sample(8000, replace=False) new_df6 = pd.concat([new_df1, new_df2, new_df3, new_df4, new_df5], axis=0, ignore_index=True) new_df6.to_csv(f"{counter}.csv", index=False) counter +=1
3. 补充:处理样本不足的情况
如果某个评分的样本数不够4000/8000,sample()会报错,建议加个安全采样的小函数:
def safe_sample(df, target_num): if len(df) >= target_num: return df.sample(target_num, replace=False) else: print(f"警告:{target_num}条样本不足,仅{len(df)}条,将重复采样补足") return df.sample(target_num, replace=True) # 替换原sample调用 new_df1 = safe_sample(new_df[new_df['overall'] ==1], 4000)
4. 合并CSV的优化(可选)
合并CSV时可以直接用生成器表达式,省内存:
from glob import glob filenames = glob('*.csv') finaldf = pd.concat([pd.read_csv(f) for f in filenames], axis=0, ignore_index=True) finaldf.to_csv("balanced_reviews.csv", index=False)
按上面的步骤修改后,应该就能顺利生成平衡的评论数据集了。
内容的提问来源于stack exchange,提问作者khushi khandelwal
相关产品推荐
相关产品推荐

