You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取多个JSON文件信息并转换为pandas DataFrame

现有代码问题
  • 循环内data变量每次被新读取的JSON内容覆盖,循环结束后仅保留最后一个文件的内容
  • json.loads()用于解析字符串格式的JSON,直接传入文件对象会触发类型错误,需改用json.load()读取文件对象
提取信息并转换为Pandas DataFrame 实现方案

核心逻辑为:初始化空列表存储所有提取结果,遍历读取每个JSON文件后提取目标字段存入列表,最终将列表转换为DataFrame。
示例代码如下:

import os
import json
import pandas as pd

# 你已有的JSON文件筛选逻辑
jsons_data = json_import(path_to_json)
# 初始化空列表存储所有提取记录
extracted_records = []

for json_file_name in jsons_data:
    full_file_path = os.path.join(path_to_json, json_file_name)
    with open(full_file_path, 'r', encoding='utf-8') as f:
        json_data = json.load(f)
    
    # 按需替换为你要提取的字段,支持嵌套字段取值
    current_record = {
        "源文件名": json_file_name,
        "目标字段1": json_data.get("目标字段1"),
        "目标字段2": json_data.get("目标字段2"),
        # 嵌套字段取值示例:json_data.get("一级字段", {}).get("二级字段")
    }
    extracted_records.append(current_record)

# 转换为DataFrame
result_df = pd.DataFrame(extracted_records)
# 查看前5条结果验证
print(result_df.head())
可选扩展操作
  • 导出为CSV文件:result_df.to_csv("json提取结果.csv", index=False, encoding='utf-8-sig')
  • 导出为Excel文件:result_df.to_excel("json提取结果.xlsx", index=False)
  • 过滤空值:result_df = result_df.dropna(subset=["必填字段名"]) 可删除缺少必填字段的记录

内容的提问来源于stack exchange,提问作者Felipe Roque

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 17:45:02