You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中按分组筛选指定后缀文件名并传入函数的方法

问题描述

我有一个文件名列表clist,示例片段如下:

clist[0:31]

['ACCESS-CM2.ssp245.r3i1p1f1.2045-2074.LOCA_16thdeg_exceedence_summary_2.0sd.csv',
'ACCESS-CM2.ssp245.r3i1p1f1.2045-2074.LOCA_16thdeg_pr_maxexceed.nc', 
...
'ACCESS-CM2.ssp370.r1i1p1f1.2045-2074.LOCA_16thdeg_tmin_exceedlt.nc']

需要按模型、SSP和realization的唯一组合对文件分组,从每组中筛选出以"tavg_monmean"和"pr_monsum"结尾的文件,传入以下函数处理:

def little_function(tavg_monmean_input, pr_monsum_input):
    <do stuff with inputs>

我尝试用嵌套循环实现但没成功,示例代码如下:

import re

# Create short list of unique model, SSP, realization
uqlis = [re.search("[A-Z].*\.[0-9]{4}", i).group(0) for i in clist]
unqlis = set(uqlis)
unqlis = list(unqlis)


for i in range(len(unqlis)):
    # subset of clist that matches one of the unique model, etc
    clsub = [clist[li] for li in range(len(clist)) if unqlis[i] in clist[li]]
    for j in range(len(clsub)):
        ###1 Select the file from clsub that ends with tavg_monmean for
        ### the current group
        if "tavg_monmean" in clsub[j]:
            tam = clsub[j]
        ###2 Select the file from clsub that ends with pr_monsum for the
        ### current group
        elif "pr_monsum" in clsub[j]:
            psm = clsub[j]
        print(tam)
        print(psm)
        # provide the tam and psm (the tavg_monmean and pr_monsum variables) to the function 
        little_function(tam, psm)

请问如何正确实现按组筛选指定后缀文件并传入函数的逻辑?

解决方案

你的代码核心问题有两个:一是内层循环遍历单个文件就直接调用函数,此时可能仅找到一个目标文件,变量未完全赋值会报错;二是分组用的正则表达式匹配范围不准确,容易导致分组错误。以下是修正后的可靠实现:

步骤1:精准提取分组键

文件名格式为模型.SSP.realization.其他内容,直接按.拆分前三个字段作为分组标识,比正则更精准:

def get_group_key(filename):
    # 拆分文件名,取前三个字段作为分组键:模型、SSP、realization
    parts = filename.split('.')
    return tuple(parts[:3])

步骤2:用字典归类分组文件

借助defaultdict把同一组的文件归类,键为分组标识,值为该组的文件列表:

from collections import defaultdict

# 初始化分组字典
grouped_files = defaultdict(list)
for file in clist:
    key = get_group_key(file)
    grouped_files[key].append(file)

步骤3:遍历分组筛选文件并调用函数

对每个分组,先遍历所有文件找到两个目标后缀的文件,确认都找到后再调用函数,避免因缺失文件引发错误:

for group_key, files in grouped_files.items():
    tavg_file = None
    pr_file = None
    
    # 遍历组内文件,筛选目标
    for file in files:
        if file.endswith('tavg_monmean'):
            tavg_file = file
        elif file.endswith('pr_monsum'):
            pr_file = file
    
    # 确认两个文件都存在才执行函数
    if tavg_file and pr_file:
        little_function(tavg_file, pr_file)
    else:
        # 可选:处理缺少文件的情况
        print(f"分组 {group_key} 缺少目标文件:tavg_monmean={bool(tavg_file)}, pr_monsum={bool(pr_file)}")

完整整合代码

from collections import defaultdict

def get_group_key(filename):
    parts = filename.split('.')
    return tuple(parts[:3])

def little_function(tavg_monmean_input, pr_monsum_input):
    # 替换为你的业务逻辑
    print(f"处理tavg文件:{tavg_monmean_input}")
    print(f"处理pr文件:{pr_monsum_input}")

# 示例测试用clist
clist = [
    'ACCESS-CM2.ssp245.r3i1p1f1.2045-2074.LOCA_16thdeg_tavg_monmean',
    'ACCESS-CM2.ssp245.r3i1p1f1.2045-2074.LOCA_16thdeg_pr_monsum',
    'ACCESS-CM2.ssp370.r1i1p1f1.2045-2074.LOCA_16thdeg_tavg_monmean',
    'ACCESS-CM2.ssp370.r1i1p1f1.2045-2074.LOCA_16thdeg_pr_monsum',
    'ACCESS-CM2.ssp245.r3i1p1f1.2045-2074.LOCA_16thdeg_exceedence_summary_2.0sd.csv'
]

# 执行分组与处理逻辑
grouped_files = defaultdict(list)
for file in clist:
    key = get_group_key(file)
    grouped_files[key].append(file)

for group_key, files in grouped_files.items():
    tavg_file = None
    pr_file = None
    for file in files:
        if file.endswith('tavg_monmean'):
            tavg_file = file
        elif file.endswith('pr_monsum'):
            pr_file = file
    if tavg_file and pr_file:
        little_function(tavg_file, pr_file)
    else:
        print(f"分组 {group_key} 缺少必要文件,跳过处理")

原代码失败原因总结

  1. 内层循环每次仅检查一个文件就立刻调用函数,大概率只有一个变量被赋值,未定义的变量会触发NameError
  2. 正则表达式[A-Z].*\.[0-9]{4}会匹配到文件名中多余内容(比如包含2045-2074前缀),导致分组不准确
  3. 未处理分组中缺少目标文件的场景,容错性差

内容的提问来源于stack exchange,提问作者John Polo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.02 07:20:33