You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量读取JSON文件时如何忽略缺失指定列的文件以规避KeyError

解决方案

完全可以通过try-except捕获KeyError实现异常文件的跳过,同时修复原代码中读取文件未拼接路径的潜在问题,修改后代码如下:

import os, json
import pandas as pd

path_to_json = 'C:/Users/aaa/Desktop/'
json_files = [pos_json for pos_json in os.listdir(path_to_json) if pos_json.endswith('.json')]

def func(s):
    try:
        return eval(s)
    except:
        return dict()

list_of_df=[]
required_cols = ["xxx", "aaa"]
for file_name in json_files:
    # 拼接完整文件路径,避免工作目录不对导致读不到文件
    full_path = os.path.join(path_to_json, file_name)
    try:
        df = pd.read_json(full_path, lines=True)
        df = df[['columnx']]
        df = df['columnx'].apply(func)
        df = pd.json_normalize(df)
        df = pd.DataFrame(df[required_cols])
        list_of_df.append(df)
    except KeyError as e:
        print(f"文件{file_name}缺失必要列,错误信息:{e},已跳过")
        continue
    except Exception as e:
        print(f"文件{file_name}解析失败,错误信息:{e},已跳过")
        continue

# 先判断有没有有效数据,避免concat报错
if list_of_df:
    df = pd.concat(list_of_df, ignore_index=True)
    # 如果需要Index列是pandas行索引,执行下一行即可
    # df = df.reset_index(names='Index')
    print(df[['xxx', 'aaa']].head())
else:
    print("没有符合要求的JSON文件")

实现说明

  • 所有单文件处理逻辑都包裹在try块内,只要触发列缺失的KeyError、JSON解析错误等异常,就会进入对应异常分支跳过当前文件,继续处理下一个
  • 也可在取列前增加if all(col in df.columns for col in required_cols)的预判断逻辑,和异常捕获的效果完全一致
  • 修复了原代码只传文件名未拼接路径的问题,避免运行时找不到JSON文件的错误
  • 增加了空列表判断逻辑,避免所有文件都不合格时pd.concat触发报错
  • 如果你的Index是数据本身的字段,直接加入required_cols列表即可

内容的提问来源于stack exchange,提问作者xavi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 21:45:00