You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取多JSON文件时遭遇JSONDecodeError及MemoryError求助

解决JSON文件读取的JSONDecodeError与MemoryError问题

一、先解决JSONDecodeError问题

报错JSONDecodeError: Expecting value: line 1 column 1 (char 0),本质是文件无法被解析为有效JSON,常见原因和修复方式如下:

常见原因

  • 存在空JSON文件(大小为0字节)
  • 文件编码非UTF-8,读取时未指定编码导致乱码
  • 文件内容不是标准JSON格式(比如语法错误、多行JSON对象但用了整读方式)
  • 混入了非JSON文件(比如系统隐藏文件)

修复方案

  1. 添加异常捕获与无效文件排查
    修改load_json_to_dataframe函数,捕获解析错误并记录问题文件名,避免整个流程中断:
import os
import pandas as pd
import json

def load_json_to_dataframe(json_file_path):
    try:
        # 跳过空文件
        if os.path.getsize(json_file_path) == 0:
            print(f"跳过空文件: {json_file_path}")
            return None
        # 指定UTF-8编码读取,避免编码问题
        with open(json_file_path, 'r', encoding='utf-8') as json_file:
            # 先尝试整读解析
            try:
                doc = json.load(json_file)
                return pd.json_normalize(doc)
            # 如果是多行JSON(每行一个对象),改用逐行读取
            except json.JSONDecodeError:
                json_file.seek(0)
                data = [json.loads(line) for line in json_file if line.strip()]
                return pd.json_normalize(data)
    except Exception as e:
        print(f"处理文件失败 {json_file_path}: {str(e)}")
        return None
  1. 精准匹配JSON文件
    修改read_json_files函数,用glob更精准过滤JSON文件,同时跳过加载失败的返回值:
def read_json_files(folder_path):
    dataframes = []
    import glob
    json_files = glob.glob(os.path.join(folder_path, "*.json"))
    
    for json_file in json_files:
        df = load_json_to_dataframe(json_file)
        if df is not None and not df.empty:
            dataframes.append(df)
    if not dataframes:
        print("没有有效数据可合并")
        return pd.DataFrame()
    return pd.concat(dataframes, ignore_index=True)

二、解决后续的MemoryError问题

删除无效文件后出现MemoryError,是因为加载的所有DataFrame都存在内存中,合并时内存不足。修复思路是减少内存占用或避免全量内存存储:

方案1:分批写入磁盘,避免全量内存存储

不要把所有DataFrame存在列表里,而是每处理一个就追加到磁盘文件(Parquet格式比CSV更节省空间):

def read_json_files_to_disk(folder_path, output_path):
    import glob
    json_files = glob.glob(os.path.join(folder_path, "*.json"))
    first_file = True
    
    for json_file in json_files:
        df = load_json_to_dataframe(json_file)
        if df is not None and not df.empty:
            # 第一个文件写入时覆盖,后续追加
            df.to_parquet(output_path, mode='w' if first_file else 'append', 
                          compression='snappy')
            first_file = False
    # 最后读取合并后的文件
    return pd.read_parquet(output_path)

# 使用示例
folder_path = 'C:/Users/gusta/Desktop/business/Emprendimiento'
combined_dataframe = read_json_files_to_disk(folder_path, 'combined_data.parquet')

方案2:优化数据类型减少内存占用

加载DataFrame后,主动优化列的数据类型,压缩内存占用:

def optimize_df_memory(df):
    # 优化数值列,向下转换类型
    for col in df.select_dtypes(include=['int64', 'float64']).columns:
        df[col] = pd.to_numeric(df[col], downcast='integer' if df[col].dtype == 'int64' else 'float')
    # 字符串列去重率高的转为category类型
    for col in df.select_dtypes(include=['object']).columns:
        if df[col].nunique() / len(df) < 0.5:
            df[col] = df[col].astype('category')
    return df

# 在load函数中加入优化
def load_json_to_dataframe(json_file_path):
    # ... 原有代码 ...
    if df is not None:
        df = optimize_df_memory(df)
    return df

内容的提问来源于stack exchange,提问作者Héctor Garrido

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 05:22:44