You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Modin与Ray读取Excel后无法移动文件,求解决方法

解决Modin+Ray读取Excel后文件被占用无法移动的问题

问题核心在于Modin基于Ray分布式执行时,pd.read_excel调用后,Ray的worker进程可能仍持有原Excel文件的句柄,导致本地进程执行shutil.move时提示文件被占用。以下是两种可行的解决方案:

方案1:通过内存中转读取文件,彻底释放原文件句柄

先将原文件读取到内存中的BytesIO对象,再让Modin从内存读取数据,这样原文件在读取完成后会被立即关闭,不会被进程占用:

import os
from pathlib import Path
import shutil
import ray
import io
ray.init()
import modin.pandas as pd

current_directory = os.getcwd()
import_folder_path = os.path.join(current_directory, 'IMPORT')
folder_path: Path = Path(import_folder_path)
file_list = [f for f in os.listdir(folder_path) if f.endswith('.xlsx')]

df2 = pd.DataFrame()
if file_list:
    imported_file_path = os.path.join(current_directory, 'IMPORTED')
    # 确保目标文件夹存在,不存在则自动创建
    os.makedirs(imported_file_path, exist_ok=True)
    
    for file in file_list:
        file_path = os.path.join(folder_path, file)
        # 读取文件到内存缓冲区,with块结束后自动关闭原文件
        with open(file_path, 'rb') as f:
            file_buffer = io.BytesIO(f.read())
        
        # 从内存缓冲区读取Excel数据
        df = pd.read_excel(file_buffer)
        df = df[df['Delivery Status'] != 'Delivered']
        # 正确累加数据(原代码的df.append(df)会重复添加当前df)
        df2 = pd.concat([df2, df], ignore_index=True)
        
        # 此时原文件已关闭,可正常执行移动操作
        shutil.move(file_path, os.path.join(imported_file_path, file))

    output_file_path = os.path.join(folder_path, 'output.xlsx')
    df2.to_excel(output_file_path, index=False)
else:
    print("No excel file found")

方案2:显式释放资源后再移动文件

如果不想修改读取逻辑,可以在处理完单文件数据后,显式删除DataFrame引用、触发垃圾回收,并强制Modin释放Ray集群中的相关资源:

# 仅修改原代码的循环部分
for file in file_list:
    file_path = os.path.join(folder_path,file)
    df = pd.read_excel(file_path)
    df = df[df['Delivery Status'] != 'Delivered']
    df2 = pd.concat([df2, df], ignore_index=True)
    
    # 释放当前文件关联的资源
    del df
    import gc
    gc.collect()
    # 强制Modin释放Ray集群中的IO相关资源
    from modin.core.execution.ray.implementations.pandas_on_ray.io import PandasOnRayIO
    PandasOnRayIO._release_resources()
    
    shutil.move(file_path, os.path.join(imported_file_path, file))

额外注意事项

  • 原代码中df2 = df.append(df)会将当前df重复添加两次,建议改用pd.concat实现正确的数据累加
  • 路径拼接统一使用os.path.join,避免跨平台路径格式问题
  • 提前创建IMPORTED文件夹,防止移动时因目标路径不存在报错

内容的提问来源于stack exchange,提问作者Angelo Malfitano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 00:11:04