You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

排查pandas DataFrame.to_json()百万级数据写入失败原因

排查pandas DataFrame.to_json()写入百万级数据失败的问题

针对你遇到的百万级DataFrame调用to_json()写入失败的问题,结合你的代码和场景,我整理了几个常见原因和对应的解决方案:

一、先排查基础环境问题

  • 磁盘空间不足:百万条记录的JSON文件体积不小,先确认目标磁盘还有足够剩余空间。Linux/macOS可以用df -h命令查看,Windows直接看资源管理器的磁盘容量。
  • 文件权限问题:确保你当前运行程序的用户对write_file指定的路径有写入权限,比如Linux下用ls -l检查目标目录的权限设置。
  • 系统内存限制:一次性处理百万条数据可能耗尽内存,导致写入中断,这种情况可能会触发MemoryError,但你的代码目前没捕获这个异常。

二、针对大数据量的优化方案

1. 分块写入DataFrame

一次性写入百万条数据会占用大量内存,分块写入可以显著降低内存压力,同时避免写入中断:

import pandas as pd
import sys

chunk_size = 10000  # 可根据你的内存情况调整块大小
write_file = "your_target_file.json"

try:
    with open(write_file, 'w') as f:
        # 按块遍历DataFrame
        for idx in range(0, len(predictions), chunk_size):
            chunk = predictions.iloc[idx:idx+chunk_size]
            # 第一块直接写入,后续块添加换行符保证格式正确
            if idx == 0:
                f.write(chunk.to_json(orient='records', lines=True))
            else:
                f.write('\n' + chunk.to_json(orient='records', lines=True))
except IOError as ioerr:
    print(f"IOError详情: {ioerr}")
    sys.exit(f'\n无法写入文件 ({write_file})! IOError. 退出...')
except MemoryError:
    print("写入JSON时内存不足!")
    sys.exit(f'\n无法写入文件 ({write_file})! MemoryError. 退出...')
except Exception as e:
    print(f"未预期错误: {type(e).__name__} - {e}")
    sys.exit(f'\n无法写入文件 ({write_file})! 未预期错误. 退出...')

2. 预处理特殊数据类型

虽然你说数据有效,但JSON不支持NaN、Infinity这类特殊值,可能会导致写入失败。建议先处理这些值:

# 将NaN/NA替换为None,Infinity替换为字符串(或你需要的其他值)
predictions = predictions.replace(
    {pd.NA: None, float('inf'): 'Infinity', float('-inf'): '-Infinity'}
)
# 可选:控制浮点数精度,避免不必要的格式问题
predictions.to_json(write_file, orient='records', lines=True, double_precision=10)

3. 扩展异常捕获范围

你的代码目前只捕获了EOFError和IOError,但大数据量场景下更可能遇到MemoryError,建议补充捕获并打印详细错误信息,方便定位问题:

try:
    predictions.to_json(write_file, orient='records', lines=True)
except EOFError as eoferr:
    print(f"EOFError详情: {eoferr}")
    sys.exit(f'\n无法写入文件 ({write_file})! EOFError. 退出...')
except IOError as ioerr:
    print(f"IOError详情: {ioerr}")
    sys.exit(f'\n无法写入文件 ({write_file})! IOError. 退出...')
except MemoryError:
    print("写入JSON时内存耗尽!")
    sys.exit(f'\n无法写入文件 ({write_file})! MemoryError. 退出...')
except Exception as e:
    print(f"错误类型: {type(e).__name__}, 详情: {e}")
    sys.exit(f'\n无法写入文件 ({write_file})! 未知错误. 退出...')

4. 改用原生json模块逐行写入

如果分块写入还是有问题,可以用Python原生json模块配合DataFrame的迭代器逐行写入,内存占用更低:

import json
import sys

write_file = "your_target_file.json"

try:
    with open(write_file, 'w') as f:
        # to_dict('records')在pandas 1.5+版本返回迭代器,节省内存
        for record in predictions.to_dict('records'):
            json.dump(record, f)
            f.write('\n')
except Exception as e:
    print(f"错误详情: {e}")
    sys.exit(f'\n无法写入文件 ({write_file})! 退出...')

内容的提问来源于stack exchange,提问作者tatlar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:48:50