You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量读取.msg文件时出现List index out of range错误的求助

批量读取.msg文件时List index out of range错误排查与修复

问题背景

批量读取.msg雨量报告文件时触发List index out of range错误,目标是提取数据生成DataFrame:每行对应一个雨量计,列名为「时间戳_PastHour」,单元格存储对应时段的降雨量。

错误信息

Error processing file FW_ Rain Report at 7_2_2020 10_39_40 PM.msg: list index out of range
Error processing file FW_ Rain Report at 7_2_2020 11_39_40 PM.msg: list index out of range

原始代码

import os
import extract_msg
import pandas as pd
from datetime import datetime

# Directory containing .msg files
directory = 'C:\\Users'

# Initialize an empty dictionary to store the data
data = {}

# Loop through each file in the directory
for filename in os.listdir(directory):
    if filename.endswith('.msg'):
        try:
            # Extract the content of the .msg file
            msg = extract_msg.Message(os.path.join(directory, filename))
            content = msg.body
            
            # Extract the timestamp
            timestamp_line = content.split('\n')[3]
            timestamp_str = timestamp_line.split('at ')[1].strip()
            timestamp = datetime.strptime(timestamp_str, '%m/%d/%Y %I:%M:%S %p')
            
            # Extract the Rain Gauge data
            lines = content.split('\n')
            for line in lines:
                if line and line.split()[0] != 'Gauge':
                    parts = line.split()
                    gauge = parts[0]
                    past_hour = parts[1]
                    
                    # Add the data to the dictionary
                    if gauge not in data:
                        data[gauge] = {}
                    data[gauge][timestamp] = past_hour
        except Exception as e:
            print(f"Error processing file {filename}: {e}")

# Convert the dictionary to a DataFrame
df = pd.DataFrame(data).T

# Sort the columns (timestamps)
df = df.sort_index(axis=1)

# Save the DataFrame to a CSV file
df.to_csv('rain_gauge_data.csv')

print("Data has been saved to 'rain_gauge_data.csv'")

.msg文件内容示例


发件人: Random user@gmail.com
发送时间: 2020年7月2日星期日 晚上9:39:41 (UTC-05:00) 美国东部时间(美国和加拿大)
收件人: Gauge-Rain Notification
主题: Rain Report at 7/2/2020 9:39:40 PM

Rain Report at 7/2/2020 9:39:40 PM

Rain Past This
Gauge Hour Event**

Pereira 0.00 0.00
Puebla 0.30 0.49
CI 0.00 0.11
Tokito 0.01 0.18
CO 0.00 0.04
KP N/A N/A
DSS 0.00 0.00
PL 0.00 0.00
TSM 0.00 0.00
PKP 0.00 0.01
RP 1.00 1.42
HP 0.00 0.00
GG 0.20 0.45
BB 0.00 0.28

**自本次降雨事件开始以来(可能包含午夜前的降雨量)。每个Rain Gauge有独立的降雨事件。

Flow Report at 7/2/2020 9:39:40 PM

错误原因与修复方案

错误根源

  1. 硬编码行号提取时间戳:依赖content.split('\n')[3]获取时间戳行,转发文件(FW_开头)的邮件头行数可能与示例不一致,导致索引越界。
  2. 未校验行分割结果:直接访问parts[0]和parts[1],若行内容为空或分割后元素不足2个,触发索引越界。
  3. 未限定数据行范围:遍历所有行时会处理非雨量计数据(如底部的Flow Report行),引发错误。

修复后的代码

import os
import extract_msg
import pandas as pd
from datetime import datetime

directory = 'C:\\Users'
data = {}

for filename in os.listdir(directory):
    if filename.endswith('.msg'):
        try:
            msg = extract_msg.Message(os.path.join(directory, filename))
            content = msg.body
            # 过滤空行并去除首尾空格
            lines = [line.strip() for line in content.split('\n') if line.strip()]
            
            # 动态提取时间戳:从Rain Report行获取
            timestamp = None
            for line in lines:
                if line.startswith('Rain Report at'):
                    timestamp_str = line.split('at ')[1].strip()
                    timestamp = datetime.strptime(timestamp_str, '%m/%d/%Y %I:%M:%S %p')
                    break
            if not timestamp:
                print(f"警告:{filename}中未找到时间戳")
                continue
            
            # 定位雨量计数据的起止行
            data_start = None
            data_end = None
            for idx, line in enumerate(lines):
                if line in ('Rain    Past    This', 'Gauge   Hour    Event**'):
                    data_start = idx + 1
                elif data_start is not None and (line.startswith('**') or line.startswith('Flow Report')):
                    data_end = idx
                    break
            
            if not data_start:
                print(f"警告:{filename}中未找到雨量计数据")
                continue
            
            # 处理有效数据行
            for line in lines[data_start:data_end]:
                parts = line.split()
                if len(parts) >= 2:
                    gauge = parts[0]
                    past_hour = parts[1]
                    if gauge not in data:
                        data[gauge] = {}
                    data[gauge][timestamp] = past_hour
                else:
                    print(f"警告:{filename}中存在无效数据行:{line}")
                    
        except Exception as e:
            print(f"处理文件{filename}出错:{e}")

# 生成DataFrame并保存
if data:
    df = pd.DataFrame(data).T
    df = df.sort_index(axis=1)
    # 重命名列为「时间戳_PastHour」格式
    df.columns = [f"{col.strftime('%Y-%m-%d %H:%M:%S')}_PastHour" for col in df.columns]
    df.to_csv('rain_gauge_data.csv')
    print("数据已保存至'rain_gauge_data.csv'")
else:
    print("未提取到有效数据")

修复说明

  • 动态提取时间戳:遍历行寻找Rain Report at开头的行,避免硬编码行号的依赖。
  • 限定数据行范围:找到表头后开始处理,遇到注释或其他报告时停止,只处理有效雨量计数据。
  • 增加索引校验:分割行后先检查元素数量,再访问索引,防止越界。
  • 列名格式化:将时间戳列重命名为需求指定的格式。

内容的提问来源于stack exchange,提问作者Jose Vasquez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 09:53:13