批量读取.msg文件时出现List index out of range错误的求助
问题背景
批量读取.msg雨量报告文件时触发List index out of range错误,目标是提取数据生成DataFrame:每行对应一个雨量计,列名为「时间戳_PastHour」,单元格存储对应时段的降雨量。
错误信息
Error processing file FW_ Rain Report at 7_2_2020 10_39_40 PM.msg: list index out of range
Error processing file FW_ Rain Report at 7_2_2020 11_39_40 PM.msg: list index out of range
原始代码
import os import extract_msg import pandas as pd from datetime import datetime # Directory containing .msg files directory = 'C:\\Users' # Initialize an empty dictionary to store the data data = {} # Loop through each file in the directory for filename in os.listdir(directory): if filename.endswith('.msg'): try: # Extract the content of the .msg file msg = extract_msg.Message(os.path.join(directory, filename)) content = msg.body # Extract the timestamp timestamp_line = content.split('\n')[3] timestamp_str = timestamp_line.split('at ')[1].strip() timestamp = datetime.strptime(timestamp_str, '%m/%d/%Y %I:%M:%S %p') # Extract the Rain Gauge data lines = content.split('\n') for line in lines: if line and line.split()[0] != 'Gauge': parts = line.split() gauge = parts[0] past_hour = parts[1] # Add the data to the dictionary if gauge not in data: data[gauge] = {} data[gauge][timestamp] = past_hour except Exception as e: print(f"Error processing file {filename}: {e}") # Convert the dictionary to a DataFrame df = pd.DataFrame(data).T # Sort the columns (timestamps) df = df.sort_index(axis=1) # Save the DataFrame to a CSV file df.to_csv('rain_gauge_data.csv') print("Data has been saved to 'rain_gauge_data.csv'")
.msg文件内容示例
发件人: Random user@gmail.com
发送时间: 2020年7月2日星期日 晚上9:39:41 (UTC-05:00) 美国东部时间(美国和加拿大)
收件人: Gauge-Rain Notification
主题: Rain Report at 7/2/2020 9:39:40 PMRain Report at 7/2/2020 9:39:40 PM
Rain Past This
Gauge Hour Event**Pereira 0.00 0.00
Puebla 0.30 0.49
CI 0.00 0.11
Tokito 0.01 0.18
CO 0.00 0.04
KP N/A N/A
DSS 0.00 0.00
PL 0.00 0.00
TSM 0.00 0.00
PKP 0.00 0.01
RP 1.00 1.42
HP 0.00 0.00
GG 0.20 0.45
BB 0.00 0.28**自本次降雨事件开始以来(可能包含午夜前的降雨量)。每个Rain Gauge有独立的降雨事件。
Flow Report at 7/2/2020 9:39:40 PM
错误原因与修复方案
错误根源
- 硬编码行号提取时间戳:依赖
content.split('\n')[3]获取时间戳行,转发文件(FW_开头)的邮件头行数可能与示例不一致,导致索引越界。 - 未校验行分割结果:直接访问
parts[0]和parts[1],若行内容为空或分割后元素不足2个,触发索引越界。 - 未限定数据行范围:遍历所有行时会处理非雨量计数据(如底部的Flow Report行),引发错误。
修复后的代码
import os import extract_msg import pandas as pd from datetime import datetime directory = 'C:\\Users' data = {} for filename in os.listdir(directory): if filename.endswith('.msg'): try: msg = extract_msg.Message(os.path.join(directory, filename)) content = msg.body # 过滤空行并去除首尾空格 lines = [line.strip() for line in content.split('\n') if line.strip()] # 动态提取时间戳:从Rain Report行获取 timestamp = None for line in lines: if line.startswith('Rain Report at'): timestamp_str = line.split('at ')[1].strip() timestamp = datetime.strptime(timestamp_str, '%m/%d/%Y %I:%M:%S %p') break if not timestamp: print(f"警告:{filename}中未找到时间戳") continue # 定位雨量计数据的起止行 data_start = None data_end = None for idx, line in enumerate(lines): if line in ('Rain Past This', 'Gauge Hour Event**'): data_start = idx + 1 elif data_start is not None and (line.startswith('**') or line.startswith('Flow Report')): data_end = idx break if not data_start: print(f"警告:{filename}中未找到雨量计数据") continue # 处理有效数据行 for line in lines[data_start:data_end]: parts = line.split() if len(parts) >= 2: gauge = parts[0] past_hour = parts[1] if gauge not in data: data[gauge] = {} data[gauge][timestamp] = past_hour else: print(f"警告:{filename}中存在无效数据行:{line}") except Exception as e: print(f"处理文件{filename}出错:{e}") # 生成DataFrame并保存 if data: df = pd.DataFrame(data).T df = df.sort_index(axis=1) # 重命名列为「时间戳_PastHour」格式 df.columns = [f"{col.strftime('%Y-%m-%d %H:%M:%S')}_PastHour" for col in df.columns] df.to_csv('rain_gauge_data.csv') print("数据已保存至'rain_gauge_data.csv'") else: print("未提取到有效数据")
修复说明
- 动态提取时间戳:遍历行寻找
Rain Report at开头的行,避免硬编码行号的依赖。 - 限定数据行范围:找到表头后开始处理,遇到注释或其他报告时停止,只处理有效雨量计数据。
- 增加索引校验:分割行后先检查元素数量,再访问索引,防止越界。
- 列名格式化:将时间戳列重命名为需求指定的格式。
内容的提问来源于stack exchange,提问作者Jose Vasquez

