os.walk遍历目录时找不到已存在的JSON文件,报FileNotFoundError
推文解析项目中FileNotFoundError问题排查与解决
问题重现
你在扩展推文特征提取后重新运行代码时遇到了FileNotFoundError,报错提示找不到1.json,但该文件明明存在于目标目录下。你的代码逻辑是遍历指定目录下的所有JSON文件,解析后提取特征生成DataFrame。
错误信息如下:
--------------------------------------------------------------------------- FileNotFoundError Traceback (most recent call last) <ipython-input-21-fba7b7321509> in <module>() 11 for file in files: 12 if file.endswith('.json'): ---> 13 with open(file, 'r') as input_file: # print (tweet.keys()) 14 for line in input_file: 15 try: FileNotFoundError: [Errno 2] No such file or directory: '1.json'
问题根源
这是一个典型的文件路径拼接错误:
os.walk()返回的file变量只是文件的纯文件名(比如1.json),而不是完整的文件路径- 当
os.walk遍历到子目录时,你的代码尝试在当前工作目录(而不是子目录)下打开这个文件,自然会找不到
解决办法
只需要修改文件打开的路径,用os.path.join()拼接当前遍历的目录路径和文件名,就能得到正确的文件绝对路径。同时还可以优化代码的健壮性:
修改后的完整代码
import os import json import pandas as pd import numpy as np from collections import defaultdict elements_keys = ['created_at', 'text', 'lang', 'geo', 'location', 'quote_count', 'reply_count', 'retweet_count', 'favorite_count', 'in_reply_to_screen_name', 'screen_name', 'description', 'verified', 'followers_count', 'friends_count', 'listed_count', 'favourites_count', 'statuses_count'] elements = defaultdict(list) # 遍历目录,注意拼接完整路径 for dirs, subdirs, files in os.walk('/Users/user/Desktop/'): for file in files: if file.endswith('.json'): # 拼接当前目录和文件名,得到完整路径 full_file_path = os.path.join(dirs, file) with open(full_file_path, 'r') as input_file: for line in input_file: try: tweet = json.loads(line) # 检查所有需要的key是否存在,不存在会抛出KeyError items = [(key, tweet[key]) for key in elements_keys] for key, value in items: elements[key].append(value) except json.JSONDecodeError: print(f"解析JSON失败:{full_file_path}的某一行格式错误") continue except KeyError as e: print(f"推文缺少必要字段:{e},文件路径:{full_file_path}") continue except Exception as e: print(f"处理文件{full_file_path}时出现未知错误:{e}") continue # 直接用elements字典创建DataFrame,无需手动逐个指定列 df = pd.DataFrame(elements) # 可以选择性地设置created_at为索引 df = df.set_index('created_at') df.to_csv('df.csv')
关键优化点
- 路径拼接:用
os.path.join(dirs, file)生成完整文件路径,确保无论遍历到哪个子目录都能正确找到文件 - 精准异常捕获:把原来的裸
except:改成捕获具体异常(JSONDecodeError、KeyError),方便定位问题,避免隐藏其他未知错误 - 简化DataFrame创建:直接用
pd.DataFrame(elements)创建DataFrame,无需手动重复列名,减少冗余代码
内容的提问来源于stack exchange,提问作者kiton
相关产品推荐
相关产品推荐

