You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过if语句筛选指定年份.tiff文件计算年蒸散发总量

按自然年累加逐月蒸散发TIFF数据的实现方案

问题背景

  • 需求:对存储逐月蒸散发数据的TIFF文件按自然年维度求和,例如累加2007年全部12个月份的对应文件,得到年度总蒸散发量
  • 现存异常:原有代码的年份筛选逻辑失效,目录下所有年份的TIFF文件都被纳入求和计算,无法得到指定年份的统计结果

原有问题代码

def pathList (d): # d is the path to the specified directory
   
   sum_array = np.zeros((2200, 2800)) # creating empty array in which to sum monthly evap. values
   nmlist = [] # creates an empty list object in which to store the names of the .tiff files
   count = 0 # creating variable to store index of files in directory

   for item in os.scandir(d): # iterating through directory contents
     
            nmlist.append(item.name) # preparing name list of .tiff files to use in "if in" statement (see below)

            tif_file = gdal.Open(pthlist[count]) # reading .tiff via gdal
            tif_band = tif_file.GetRasterBand(1) # reading first band
            tif_arr = tif_band.ReadAsArray() # converting to numpy array
            
            if "2007" in nmlist[count]: # does the name of the file contain a specific year (e.g. 2007)?
                sum_array = sum_array + tif_arr # summing monthly evap. values
       
            count += 1 # tracking index of files in directory

   return sum_array

核心错误点

  1. 存在未定义变量问题:代码仅初始化了存储文件名的nmlist,从未定义存储文件完整路径的pthlist,运行时会直接触发名称错误,无法正常读取文件
  2. 执行逻辑顺序倒置:无论文件是否属于目标年份,都会先执行GDAL读取、数组转换的重IO操作,既浪费性能,也容易引入非目标数据的计算误差
  3. 年份匹配逻辑鲁棒性差:仅用"2007" in 文件名的模糊匹配规则,容易误匹配文件名其他位置(如版本号、处理批次号)带"2007"字符的非目标文件
  4. 存在内存泄漏风险:读取GDAL数据集后未手动释放资源,批量处理大量文件时容易占满内存

修正后可运行代码

import os
import re
import numpy as np
from osgeo import gdal

def calc_yearly_et(dir_path, target_year="2007"):
    """
    计算指定年份的累计蒸散发量
    :param dir_path: 逐月TIFF文件存储的目录路径
    :param target_year: 需要统计的目标年份,默认值为2007
    :return: 年度累计蒸散发numpy数组
    """
    # 初始化与TIFF尺寸一致的零矩阵用于累加,指定float32类型节省内存
    sum_array = np.zeros((2200, 2800), dtype=np.float32)
    # 编译正则匹配规则:确保匹配的是独立4位年份,避免误匹配其他位置的相同数字
    year_match_rule = re.compile(rf'(?<!\d){target_year}(?!\d)')

    for entry in os.scandir(dir_path):
        # 跳过子目录、非TIFF格式文件
        if not entry.is_file() or not entry.name.lower().endswith((".tif", ".tiff")):
            continue
        # 先做年份校验,不匹配的文件直接跳过,不执行后续读取操作
        if not year_match_rule.search(entry.name):
            continue
        
        # 仅对符合要求的目标年份文件执行读取、累加
        ds = gdal.Open(entry.path)
        band = ds.GetRasterBand(1)
        arr = band.ReadAsArray()
        sum_array += arr
        # 释放GDAL数据集占用的内存
        ds = None

    return sum_array

关键调整说明

  • 移除冗余的索引计数器、未定义的路径列表变量,直接调用os.scandir返回对象自带的path属性获取文件完整路径,从根源避免路径引用错误
  • 重构执行流程:按照「过滤非文件→过滤非TIFF格式→过滤非目标年份」的顺序提前筛除无效内容,仅对符合要求的文件执行IO读取,计算效率提升明显
  • 替换模糊字符串匹配为正则精确匹配:规则要求目标年份前后不能紧邻其他数字,确保匹配到的是作为独立字段的年份值,大幅降低误匹配概率
  • 增加GDAL资源释放、数组数据类型指定逻辑,避免批量处理时的内存溢出、数值精度问题
  • 将目标年份设置为可传入参数,无需修改代码内部逻辑即可切换统计其他年份的累计值

内容的提问来源于stack exchange,提问作者Tim Kerremans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.01 03:57:32