Python读取多JSON文件时遭遇JSONDecodeError及MemoryError求助
解决JSON文件读取的JSONDecodeError与MemoryError问题
一、先解决JSONDecodeError问题
报错JSONDecodeError: Expecting value: line 1 column 1 (char 0),本质是文件无法被解析为有效JSON,常见原因和修复方式如下:
常见原因
- 存在空JSON文件(大小为0字节)
- 文件编码非UTF-8,读取时未指定编码导致乱码
- 文件内容不是标准JSON格式(比如语法错误、多行JSON对象但用了整读方式)
- 混入了非JSON文件(比如系统隐藏文件)
修复方案
- 添加异常捕获与无效文件排查
修改load_json_to_dataframe函数,捕获解析错误并记录问题文件名,避免整个流程中断:
import os import pandas as pd import json def load_json_to_dataframe(json_file_path): try: # 跳过空文件 if os.path.getsize(json_file_path) == 0: print(f"跳过空文件: {json_file_path}") return None # 指定UTF-8编码读取,避免编码问题 with open(json_file_path, 'r', encoding='utf-8') as json_file: # 先尝试整读解析 try: doc = json.load(json_file) return pd.json_normalize(doc) # 如果是多行JSON(每行一个对象),改用逐行读取 except json.JSONDecodeError: json_file.seek(0) data = [json.loads(line) for line in json_file if line.strip()] return pd.json_normalize(data) except Exception as e: print(f"处理文件失败 {json_file_path}: {str(e)}") return None
- 精准匹配JSON文件
修改read_json_files函数,用glob更精准过滤JSON文件,同时跳过加载失败的返回值:
def read_json_files(folder_path): dataframes = [] import glob json_files = glob.glob(os.path.join(folder_path, "*.json")) for json_file in json_files: df = load_json_to_dataframe(json_file) if df is not None and not df.empty: dataframes.append(df) if not dataframes: print("没有有效数据可合并") return pd.DataFrame() return pd.concat(dataframes, ignore_index=True)
二、解决后续的MemoryError问题
删除无效文件后出现MemoryError,是因为加载的所有DataFrame都存在内存中,合并时内存不足。修复思路是减少内存占用或避免全量内存存储:
方案1:分批写入磁盘,避免全量内存存储
不要把所有DataFrame存在列表里,而是每处理一个就追加到磁盘文件(Parquet格式比CSV更节省空间):
def read_json_files_to_disk(folder_path, output_path): import glob json_files = glob.glob(os.path.join(folder_path, "*.json")) first_file = True for json_file in json_files: df = load_json_to_dataframe(json_file) if df is not None and not df.empty: # 第一个文件写入时覆盖,后续追加 df.to_parquet(output_path, mode='w' if first_file else 'append', compression='snappy') first_file = False # 最后读取合并后的文件 return pd.read_parquet(output_path) # 使用示例 folder_path = 'C:/Users/gusta/Desktop/business/Emprendimiento' combined_dataframe = read_json_files_to_disk(folder_path, 'combined_data.parquet')
方案2:优化数据类型减少内存占用
加载DataFrame后,主动优化列的数据类型,压缩内存占用:
def optimize_df_memory(df): # 优化数值列,向下转换类型 for col in df.select_dtypes(include=['int64', 'float64']).columns: df[col] = pd.to_numeric(df[col], downcast='integer' if df[col].dtype == 'int64' else 'float') # 字符串列去重率高的转为category类型 for col in df.select_dtypes(include=['object']).columns: if df[col].nunique() / len(df) < 0.5: df[col] = df[col].astype('category') return df # 在load函数中加入优化 def load_json_to_dataframe(json_file_path): # ... 原有代码 ... if df is not None: df = optimize_df_memory(df) return df
内容的提问来源于stack exchange,提问作者Héctor Garrido
相关产品推荐
相关产品推荐

