如何获取连续低温期最后一天的正确索引?Python气候数据处理求助
我正在完成一项Python气候数据处理作业,需要编写get_longest_freezing()函数,找出最长连续无0℃以上温度的天数,以及该时段最后一天的日期。当前代码里因为温度值存在重复,用list.index()会返回重复温度的首次出现索引,没法获取连续低温期末尾温度对应的正确索引,请问怎么获取连续序列中对应元素的索引?
作业要求
找出最长的连续无零上温度(即最高气温低于0℃)的天数,以及该时段的最后一天日期。
编写名为
get_longest_freezing()的函数,返回最长连续低温天数和该时段最后一天的日期。函数需接受两个参数:max_dates和max_temps。参考作业1中的函数调用示例,使用合理的变量名。将两个问题的答案整理成格式规范的单行输出。记得复用之前写过的格式化函数,不要在
get_longest_freezing()函数内部写打印语句。
数据文件示例
欧洲气候评估数据集(ECA&D),文件创建于2022年5月23日
本数据可免费使用,但需注明以下来源:Klein Tank, A.M.G. 等,2002年。《20世纪欧洲地面气温和降水日数据集》,《国际气候学杂志》,22卷,1441-1453页。
文件格式(缺失值代码:-9999):
- 01-06位:SOUID(数据源标识)
- 08-15位:DATE(日期,格式YYYYMMDD)
- 17-21位:TX(最高气温,单位0.1℃)
- 23-27位:Q_TX(TX的质量代码:0=有效;1=可疑;9=缺失)
本数据为荷兰德比尔站(STAID:162)的混合序列,由数据源100522、906260混合更新。更多信息见sources.txt和stations.txt文件。
SOUID, DATE, TX, Q_TX 100522,19010101, -24, 0 100522,19010102, -14, 0 100522,19010103, -6, 0 100522,19010104, -11, 0 100522,19010105, -20, 0 100522,19010106, -80, 0 100522,19010107, -68, 0 100522,19010108, -7, 0 100522,19010109, 44, 0 100522,19010110, 61, 0 100522,19010111, 51, 0 100522,19010112, -20, 0 100522,19010113, -20, 0 100522,19010114, 31, 0 100522,19010115, 54, 0 100522,19010116, 42, 0 100522,19010117, 65, 0 100522,19010118, 43, 0 100522,19010119, 64, 0 100522,19010120, 72, 0 100522,19010121, 100, 0 100522,19010122, 83, 0 100522,19010123, 83, 0 100522,19010124, 57, 0 100522,19010125, 75, 0 100522,19010126, 70, 0 100522,19010127, 90, 0 100522,19010128, 54, 0 100522,19010129, 28, 0 100522,19010130, 39, 0 100522,19010131, 22, 0 100522,19010201, 23, 0 100522,19010202, -3, 0 100522,19010203, 26, 0 100522,19010204, -19, 0 100522,19010205, 17, 0 100522,19010206, 1, 0 100522,19010207, 27, 0
当前代码
def read_data(text_file): # 从第20行开始读取文件 data_file = open(text_file, 'r') data_file_read = data_file.readlines()[20:] dates = [] temperatures = [] # 仅读取日期和温度数据 for line in data_file_read: split_data = line.split(',') date = split_data[1] validation = int(split_data[3]) # 处理缺失值代码-9999 if validation == 0: # 转换为整数 temperature_value = int(split_data[2]) # 除以10得到准确温度值 temperature = float(temperature_value / 10) dates.append(date) temperatures.append(temperature) data_file.close() return dates, temperatures max_dates, max_temps = read_data('DeBiltTempMax2022.txt') def get_longest_freezing(max_len_date, max_len_temp): count = 0 max_count = 0 longest_list = [] index_list = [] # 统计最长连续低温天数 for longest_temp in max_len_temp: if longest_temp < 0: count += 1 if count > max_count: max_count = count longest_list.append(longest_temp) if longest_temp >= 0: count = 0 for index in longest_list: temp_index = max_len_temp.index(index) index_list.append(temp_index) last_temp = longest_list[-1] first = longest_list[0] date1 = max_len_temp.index(last_temp) date2 = max_len_temp.index(first) len_date = max_dates[date1] print(date1) print(date2) print(max_count) print(longest_list) print(index_list) return len_date, max_count max_len_dates, max_len_temp = get_longest_freezing(max_dates, max_temps)
解决方案
核心问题是遍历温度列表时未跟踪索引,导致后续用index()查找重复值时出错。正确做法是在遍历过程中同时记录索引和温度值,直接定位连续低温段的起止位置,无需事后查找索引。
修改后的get_longest_freezing()函数:
def get_longest_freezing(max_dates, max_temps): current_streak = 0 longest_streak = 0 last_day_index = -1 # 记录最长低温段最后一天的索引 for idx, temp in enumerate(max_temps): if temp < 0: current_streak += 1 # 当前连续天数超过记录值时,更新最长天数和最后一天索引 if current_streak > longest_streak: longest_streak = current_streak last_day_index = idx else: # 温度回到0℃以上,重置当前连续天数 current_streak = 0 # 根据索引获取最后一天的日期,处理无低温段的边界情况 last_day_date = max_dates[last_day_index] if last_day_index != -1 else None return longest_streak, last_day_date
关键改进点
- 使用
enumerate()遍历max_temps,同时获取温度的**索引idx**和值temp,直接跟踪位置,避免重复值查找错误。 - 实时更新最长连续天数和对应最后一天索引,逻辑简洁高效,无需额外存储温度列表。
- 处理无低温段的边界情况,增强代码鲁棒性。
最后调用函数并格式化输出(复用之前的日期格式化函数):
longest_days, last_date = get_longest_freezing(max_dates, max_temps) # 日期格式化函数示例:将YYYYMMDD转为YYYY-MM-DD def format_date(date_str): return f"{date_str[:4]}-{date_str[4:6]}-{date_str[6:]}" print(f"最长连续低温天数:{longest_days}天,最后一天日期:{format_date(last_date)}")
内容的提问来源于stack exchange,提问作者Maja Kubara

