You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas分块读取JSON文件并生成合并CSV时遭遇ValueError错误求助

解决ValueError: Expected object or value及代码语法问题

咱们一步步拆解你遇到的问题,这个错误主要来自两个核心原因:JSON文件存在无效行,加上代码里的布尔筛选语法写错了,一起来修复:

1. 修复JSON读取时的无效行问题

当用pd.read_json(lines=True)读取大JSON文件时,如果某一行不是合法的JSON格式(比如引号不闭合、括号缺失),就会触发ValueError: Expected object or value。咱们可以自定义一个带错误处理的迭代器,跳过这些坏行:

import pandas as pd
import json

def read_json_safely(file_path, chunksize=1000000):
    with open(file_path, 'r', encoding='utf-8') as f:
        chunk = []
        for line_num, line in enumerate(f, 1):
            try:
                chunk.append(json.loads(line))
                if len(chunk) >= chunksize:
                    yield pd.DataFrame(chunk)
                    chunk = []
            except json.JSONDecodeError:
                print(f"跳过无效行 {line_num}: {line[:50]}...")
        if chunk:
            yield pd.DataFrame(chunk)

# 用自定义迭代器替代原pd.read_json的chunk读取
df_reader = read_json_safely('Clothing_Shoes_and_Jewelry.json', chunksize=1000000)

2. 修正布尔筛选的语法错误

你的代码里所有的评分筛选都写错了!比如new_df[new_df['overall' == 1]]应该是new_df[new_df['overall'] == 1]——你把比较符号放错了位置,导致程序试图用False('overall' ==1的结果)当列名,直接会触发KeyError。修正后的筛选代码:

counter = 1 
for chunk in df_reader: 
    new_df = chunk[['overall', 'reviewText','summary']].copy()  # 直接切片更简洁
    # 修正后的布尔筛选逻辑
    new_df1 = new_df[new_df['overall'] == 1].sample(4000, replace=False)
    new_df2 = new_df[new_df['overall'] == 2].sample(4000, replace=False)
    new_df3 = new_df[new_df['overall'] == 4].sample(4000, replace=False)
    new_df4 = new_df[new_df['overall'] == 5].sample(4000, replace=False)
    new_df5 = new_df[new_df['overall'] == 3].sample(8000, replace=False)
    
    new_df6 = pd.concat([new_df1, new_df2, new_df3, new_df4, new_df5], axis=0, ignore_index=True) 
    new_df6.to_csv(f"{counter}.csv", index=False) 
    counter +=1 

3. 补充:处理样本不足的情况

如果某个评分的样本数不够4000/8000,sample()会报错,建议加个安全采样的小函数:

def safe_sample(df, target_num):
    if len(df) >= target_num:
        return df.sample(target_num, replace=False)
    else:
        print(f"警告:{target_num}条样本不足,仅{len(df)}条,将重复采样补足")
        return df.sample(target_num, replace=True)

# 替换原sample调用
new_df1 = safe_sample(new_df[new_df['overall'] ==1], 4000)

4. 合并CSV的优化(可选)

合并CSV时可以直接用生成器表达式,省内存:

from glob import glob 

filenames = glob('*.csv')
finaldf = pd.concat([pd.read_csv(f) for f in filenames], axis=0, ignore_index=True)
finaldf.to_csv("balanced_reviews.csv", index=False)

按上面的步骤修改后,应该就能顺利生成平衡的评论数据集了。

内容的提问来源于stack exchange,提问作者khushi khandelwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 19:17:48