You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

os.walk遍历目录时找不到已存在的JSON文件,报FileNotFoundError

推文解析项目中FileNotFoundError问题排查与解决

问题重现

你在扩展推文特征提取后重新运行代码时遇到了FileNotFoundError,报错提示找不到1.json,但该文件明明存在于目标目录下。你的代码逻辑是遍历指定目录下的所有JSON文件,解析后提取特征生成DataFrame。

错误信息如下:

---------------------------------------------------------------------------
FileNotFoundError Traceback (most recent call last)
<ipython-input-21-fba7b7321509> in <module>()
     11 for file in files:
     12     if file.endswith('.json'):
---> 13         with open(file, 'r') as input_file: # print (tweet.keys())
     14             for line in input_file:
     15                 try:

FileNotFoundError: [Errno 2] No such file or directory: '1.json'

问题根源

这是一个典型的文件路径拼接错误:

  • os.walk()返回的file变量只是文件的纯文件名(比如1.json),而不是完整的文件路径
  • 当os.walk遍历到子目录时,你的代码尝试在当前工作目录(而不是子目录)下打开这个文件,自然会找不到

解决办法

只需要修改文件打开的路径,用os.path.join()拼接当前遍历的目录路径和文件名,就能得到正确的文件绝对路径。同时还可以优化代码的健壮性:

修改后的完整代码

import os
import json
import pandas as pd
import numpy as np
from collections import defaultdict

elements_keys = ['created_at', 'text', 'lang', 'geo', 'location', 'quote_count', 'reply_count', 'retweet_count', 'favorite_count', 'in_reply_to_screen_name', 'screen_name', 'description', 'verified', 'followers_count', 'friends_count', 'listed_count', 'favourites_count', 'statuses_count']
elements = defaultdict(list)

# 遍历目录,注意拼接完整路径
for dirs, subdirs, files in os.walk('/Users/user/Desktop/'):
    for file in files:
        if file.endswith('.json'):
            # 拼接当前目录和文件名,得到完整路径
            full_file_path = os.path.join(dirs, file)
            with open(full_file_path, 'r') as input_file:
                for line in input_file:
                    try:
                        tweet = json.loads(line)
                        # 检查所有需要的key是否存在,不存在会抛出KeyError
                        items = [(key, tweet[key]) for key in elements_keys]
                        for key, value in items:
                            elements[key].append(value)
                    except json.JSONDecodeError:
                        print(f"解析JSON失败:{full_file_path}的某一行格式错误")
                        continue
                    except KeyError as e:
                        print(f"推文缺少必要字段:{e},文件路径:{full_file_path}")
                        continue
                    except Exception as e:
                        print(f"处理文件{full_file_path}时出现未知错误:{e}")
                        continue

# 直接用elements字典创建DataFrame,无需手动逐个指定列
df = pd.DataFrame(elements)
# 可以选择性地设置created_at为索引
df = df.set_index('created_at')

df.to_csv('df.csv')

关键优化点

  • 路径拼接:用os.path.join(dirs, file)生成完整文件路径,确保无论遍历到哪个子目录都能正确找到文件
  • 精准异常捕获:把原来的裸except:改成捕获具体异常(JSONDecodeError、KeyError),方便定位问题,避免隐藏其他未知错误
  • 简化DataFrame创建:直接用pd.DataFrame(elements)创建DataFrame,无需手动重复列名,减少冗余代码

内容的提问来源于stack exchange,提问作者kiton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 11:37:49