You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中条件筛选指定年份.tiff文件求和失效问题求解

问题背景

需要对存储年度逐月蒸散量数据的.tiff文件执行年度求和计算,例如统计2007年全部12个月的蒸散量总和得到年度蒸散总量。原有代码的文件名过滤逻辑失效,指定目录下所有年份的.tiff文件都被纳入了求和范围,无法得到单一年份的统计结果。

具体问题

如何实现仅筛选特定年份(示例为2007年)对应的.tiff文件完成蒸散量求和计算?

原有存在问题的代码
def pathList (d): # d is the path to the specified directory
   
   sum_array = np.zeros((2200, 2800)) # creating empty array in which to sum monthly evap. values
   nmlist = [] # creates an empty list object in which to store the names of the .tiff files
   count = 0 # creating variable to store index of files in directory

   for item in os.scandir(d): # iterating through directory contents
     
            nmlist.append(item.name) # preparing name list of .tiff files to use in "if in" statement (see below)

            tif_file = gdal.Open(pthlist[count]) # reading .tiff via gdal
            tif_band = tif_file.GetRasterBand(1) # reading first band
            tif_arr = tif_band.ReadAsArray() # converting to numpy array
            
            if "2007" in nmlist[count]: # does the name of the file contain a specific year (e.g. 2007)?
                sum_array = sum_array + tif_arr # summing monthly evap. values
       
            count += 1 # tracking index of files in directory

   return sum_array

可对照.tiff文件命名样例规律调整匹配规则。

问题原因

原有代码存在两个核心问题:

  • 存在未定义变量错误:读取文件时调用的pthlist从未被定义,代码本身无法正常运行
  • 逻辑顺序错误:无论文件是否属于2007年,都会先被读取为numpy数组,再做年份判断,不仅浪费IO和计算资源,索引计数和列表维护的逻辑很容易出现匹配错位,导致过滤失效
修复后代码

调整逻辑顺序,先做文件名过滤再读取目标文件,移除冗余的索引和列表维护逻辑,代码如下:

import os
import numpy as np
from osgeo import gdal

def calc_yearly_evap(d, target_year="2007"):
    # 初始化和tiff尺寸匹配的零矩阵用于累加
    sum_array = np.zeros((2200, 2800), dtype=np.float32)
    file_count = 0

    for item in os.scandir(d):
        # 跳过子目录、非tiff格式文件
        if not item.is_file() or not item.name.lower().endswith((".tif", ".tiff")):
            continue
        # 先判断文件名是否匹配目标年份,不匹配直接跳过
        if target_year not in item.name:
            continue
        # 仅对匹配的目标年份文件执行读取和累加
        tif_ds = gdal.Open(item.path)
        tif_arr = tif_ds.GetRasterBand(1).ReadAsArray()
        sum_array += tif_arr
        # 释放gdal数据集占用的内存
        tif_ds = None
        file_count += 1

    print(f"已完成{target_year}年共{file_count}个逐月蒸散文件的求和计算")
    return sum_array

可选优化

如果文件名中其他位置可能出现和年份相同的数字(例如版本号、产品编号包含"2007"字符),可以将简单的字符串包含判断替换为正则匹配,精准提取4位年份字段做比对,彻底避免误匹配。

内容的提问来源于stack exchange,提问作者Tim Kerremans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 23:13:00