OSError(WinError433)排查:谷歌云端硬盘路径间歇性失效
问题描述
我正在排查一个间歇性出现的问题,相关函数每小时被调用一次。代码运行时长不稳定:有时能连续运行数天(数百次循环),有时仅运行一天(约20次循环)就抛出以下错误:
OSError: [WinError 433] A device which does not exist was specified: 'csv_files'
经追踪,问题源于以下函数间歇性返回空DataFrame,进而导致代码崩溃:
import pandas as pd def get_existing_df(symbol:str)-> pd.DataFrame: ''' check if file exists & get the the existing dataframe for a given symbol args: symbol (str): the ticker symbol return (df, last_trade): df: (DataFrame) of the last trades (empty df if no trades) ''' file_name = 'csv_files/all_trades_' + symbol + '.csv' if os.path.exists(file_name): df = pd.read_csv(file_name, index_col=0) else: df = pd.DataFrame() return df
已确认:
- 路径存在
- 文件存在
- 拥有访问权限(完整路径指向本地谷歌云端硬盘)
但该方法仍会间歇性返回空DataFrame并触发上述错误,请问问题原因是什么?如何解决?
问题原因分析
- 相对路径的不确定性:程序运行过程中,工作目录可能被其他操作(如文件操作、子进程启动)意外切换,导致
csv_files相对路径指向错误位置,触发WinError 433,此时os.path.exists返回False,函数返回空DataFrame。 - 谷歌云端硬盘挂载不稳定:本地挂载的谷歌云端硬盘易受网络波动、同步冲突、系统休眠唤醒等影响,出现间歇性断开,导致系统无法识别该设备/路径,进而触发路径不存在的错误。
- 竞态条件干扰:如果谷歌云端硬盘同步进程正在更新目标文件或目录,可能出现
os.path.exists检查时路径存在,但后续pd.read_csv读取时路径已不可用的情况,不过这种场景更可能直接抛出读取异常,而非返回空DataFrame。
解决方案
1. 替换相对路径为绝对路径
使用绝对路径避免工作目录变化带来的问题,修改后的代码如下:
import pandas as pd import os def get_existing_df(symbol:str)-> pd.DataFrame: ''' check if file exists & get the the existing dataframe for a given symbol args: symbol (str): the ticker symbol return (df, last_trade): df: (DataFrame) of the last trades (empty df if no trades) ''' # 获取当前脚本所在目录的绝对路径 script_dir = os.path.dirname(os.path.abspath(__file__)) # 拼接绝对路径 file_name = os.path.join(script_dir, 'csv_files', f'all_trades_{symbol}.csv') if os.path.exists(file_name): df = pd.read_csv(file_name, index_col=0) else: df = pd.DataFrame() return df
2. 增加路径校验与重试机制
针对云端硬盘挂载不稳定的情况,添加重试逻辑,提高代码鲁棒性:
import pandas as pd import os import time def get_existing_df(symbol:str, max_retries=3)-> pd.DataFrame: ''' check if file exists & get the the existing dataframe for a given symbol args: symbol (str): the ticker symbol max_retries (int): maximum retry attempts when path is unavailable return (df, last_trade): df: (DataFrame) of the last trades (empty df if no trades) ''' script_dir = os.path.dirname(os.path.abspath(__file__)) file_name = os.path.join(script_dir, 'csv_files', f'all_trades_{symbol}.csv') for attempt in range(max_retries): try: # 同时校验文件存在和目录可访问性 if os.path.exists(file_name) and os.access(os.path.dirname(file_name), os.R_OK): df = pd.read_csv(file_name, index_col=0) return df raise OSError("Path or file is unavailable") except OSError as e: if attempt < max_retries - 1: time.sleep(5) # 间隔5秒重试 continue # 重试失败返回空DataFrame return pd.DataFrame()
3. 优化谷歌云端硬盘同步配置
- 确保谷歌云端硬盘客户端为最新版本,修复已知同步bug。
- 减少不必要的同步文件/目录,降低同步冲突概率。
- 系统休眠唤醒后,添加自动检查挂载状态的逻辑,或手动确认云端硬盘正常挂载。
内容的提问来源于stack exchange,提问作者darren
相关产品推荐
相关产品推荐

