You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Path.glob结合正则匹配已知子串时无法找到存在的文件

问题解决思路

核心问题1:if(matching_files)始终返回True

pathlib.Path.glob()返回的是生成器对象,无论是否有匹配文件,生成器本身在布尔判断中都会被判定为True。你需要先将其转换为列表,或者直接迭代检查是否有内容。

核心问题2:glob未匹配到文件

你混淆了glob通配符和正则表达式的语法:

  • find -regex用的是正则规则,.*表示任意字符任意次数;
  • 但pathlib的glob()使用的是shell风格通配符:*才是匹配任意非路径分隔符的字符序列,.仅表示字面量的点。
    你要匹配包含指定ID的文件名,应该用f"*{re.escape(id)}*"作为glob模式,而不是正则格式的.*{id}.*。

修改后的代码

import pathlib
import re
import os

def rename_matching_ids(df, path, dataset, subset, view):
    pl_path = pathlib.Path(path)

    for id in df['id']:
        # 优化sub_path获取方式,避免重复查询DataFrame
        sub_path_row = df[df['id'] == id].iloc[0]
        sub_path = sub_path_row['file_dir']
        newName = f"1_{dataset}_{subset}_{view}_{id}"
        img_sub_path = pl_path.joinpath(sub_path)
        
        # 使用glob通配符模式,而非正则
        glob_pattern = f"*{re.escape(id)}*"
        matching_files = list(img_sub_path.glob(glob_pattern))

        if matching_files:
            for file in matching_files:
                # file已经是完整路径,无需再拼接img_sub_path
                print(file)
                if file.exists():
                    newPath = img_sub_path / newName
                    # 保留原文件后缀
                    newPath = newPath.with_suffix(file.suffix)
                    print(f"Renaming: {file.as_posix()} -> {newPath.as_posix()}")
                    # 使用pathlib的rename方法,更安全
                    file.rename(newPath)
                else:
                    print(f"Error: Cannot find file- {file.as_posix()}")
        else:
            print(f"Error: Cannot find ID {id}")

额外优化点

  • 避免重复查询DataFrame:用iloc[0]直接获取单行数据,比values[0]更直观;
  • 保留原文件后缀:原来的代码会丢失文件扩展名,添加with_suffix(file.suffix)修复;
  • 直接使用file变量:glob()返回的已经是完整的Path对象,无需再用img_sub_path / file拼接;
  • 用f-string替代字符串拼接,更简洁可读。

内容的提问来源于stack exchange,提问作者James Miller

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 09:12:47