You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python从多文件夹加载多个JSON文件到Pandas DataFrame的高效方法

将多层文件夹中的JSON文件合并为Pandas DataFrame的有效方法

核心思路

遍历data根目录下的所有子文件夹,筛选出所有.json文件,逐个读取为DataFrame后合并成一个统一的数据集。

具体实现代码

import pandas as pd
import os

# 指定根文件夹路径
root_folder = "data"
# 用于存储单个JSON文件的DataFrame
df_collection = []

# 递归遍历所有子文件夹
for current_dir, _, file_names in os.walk(root_folder):
    for file in file_names:
        # 只处理JSON格式文件
        if file.endswith(".json"):
            file_path = os.path.join(current_dir, file)
            # 根据JSON结构选择读取方式:常规JSON用默认配置,每行一个JSON对象加lines=True
            single_df = pd.read_json(file_path, lines=False)
            # 可选:添加来源文件列,方便后续数据溯源
            single_df["origin_file"] = file
            df_collection.append(single_df)

# 合并所有DataFrame
final_df = pd.concat(df_collection, ignore_index=True)

# 查看合并结果示例
print(final_df.head())

关键注意点

  • 如果你的JSON是每行一个独立JSON对象(NDJSON格式),需将pd.read_json的lines参数设为True。
  • 若不同JSON文件字段结构不一致,concat会自动为缺失字段填充NaN,后续可根据需求做清洗处理。
  • 处理超大量JSON文件时,可改用dask.dataframe分批次读取合并,避免内存溢出。

内容的提问来源于stack exchange,提问作者mbih

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 01:32:00