如何通过if语句筛选指定年份.tiff文件计算年蒸散发总量
按自然年累加逐月蒸散发TIFF数据的实现方案
问题背景
- 需求:对存储逐月蒸散发数据的TIFF文件按自然年维度求和,例如累加2007年全部12个月份的对应文件,得到年度总蒸散发量
- 现存异常:原有代码的年份筛选逻辑失效,目录下所有年份的TIFF文件都被纳入求和计算,无法得到指定年份的统计结果
原有问题代码
def pathList (d): # d is the path to the specified directory sum_array = np.zeros((2200, 2800)) # creating empty array in which to sum monthly evap. values nmlist = [] # creates an empty list object in which to store the names of the .tiff files count = 0 # creating variable to store index of files in directory for item in os.scandir(d): # iterating through directory contents nmlist.append(item.name) # preparing name list of .tiff files to use in "if in" statement (see below) tif_file = gdal.Open(pthlist[count]) # reading .tiff via gdal tif_band = tif_file.GetRasterBand(1) # reading first band tif_arr = tif_band.ReadAsArray() # converting to numpy array if "2007" in nmlist[count]: # does the name of the file contain a specific year (e.g. 2007)? sum_array = sum_array + tif_arr # summing monthly evap. values count += 1 # tracking index of files in directory return sum_array
核心错误点
- 存在未定义变量问题:代码仅初始化了存储文件名的
nmlist,从未定义存储文件完整路径的pthlist,运行时会直接触发名称错误,无法正常读取文件 - 执行逻辑顺序倒置:无论文件是否属于目标年份,都会先执行GDAL读取、数组转换的重IO操作,既浪费性能,也容易引入非目标数据的计算误差
- 年份匹配逻辑鲁棒性差:仅用
"2007" in 文件名的模糊匹配规则,容易误匹配文件名其他位置(如版本号、处理批次号)带"2007"字符的非目标文件 - 存在内存泄漏风险:读取GDAL数据集后未手动释放资源,批量处理大量文件时容易占满内存
修正后可运行代码
import os import re import numpy as np from osgeo import gdal def calc_yearly_et(dir_path, target_year="2007"): """ 计算指定年份的累计蒸散发量 :param dir_path: 逐月TIFF文件存储的目录路径 :param target_year: 需要统计的目标年份,默认值为2007 :return: 年度累计蒸散发numpy数组 """ # 初始化与TIFF尺寸一致的零矩阵用于累加,指定float32类型节省内存 sum_array = np.zeros((2200, 2800), dtype=np.float32) # 编译正则匹配规则:确保匹配的是独立4位年份,避免误匹配其他位置的相同数字 year_match_rule = re.compile(rf'(?<!\d){target_year}(?!\d)') for entry in os.scandir(dir_path): # 跳过子目录、非TIFF格式文件 if not entry.is_file() or not entry.name.lower().endswith((".tif", ".tiff")): continue # 先做年份校验,不匹配的文件直接跳过,不执行后续读取操作 if not year_match_rule.search(entry.name): continue # 仅对符合要求的目标年份文件执行读取、累加 ds = gdal.Open(entry.path) band = ds.GetRasterBand(1) arr = band.ReadAsArray() sum_array += arr # 释放GDAL数据集占用的内存 ds = None return sum_array
关键调整说明
- 移除冗余的索引计数器、未定义的路径列表变量,直接调用
os.scandir返回对象自带的path属性获取文件完整路径,从根源避免路径引用错误 - 重构执行流程:按照「过滤非文件→过滤非TIFF格式→过滤非目标年份」的顺序提前筛除无效内容,仅对符合要求的文件执行IO读取,计算效率提升明显
- 替换模糊字符串匹配为正则精确匹配:规则要求目标年份前后不能紧邻其他数字,确保匹配到的是作为独立字段的年份值,大幅降低误匹配概率
- 增加GDAL资源释放、数组数据类型指定逻辑,避免批量处理时的内存溢出、数值精度问题
- 将目标年份设置为可传入参数,无需修改代码内部逻辑即可切换统计其他年份的累计值
内容的提问来源于stack exchange,提问作者Tim Kerremans
相关产品推荐
相关产品推荐

