You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Gmail API获取并解析邮件时遇到两类问题求助

解决Gmail API邮件解析的两个问题

问题1:部分邮件body字段为空,BeautifulSoup解析结果含多余符号

原因分析

  • Gmail API返回的邮件分单部分和多结构(payload.parts)两种格式,不少邮件的正文实际嵌套在parts层级里,直接取顶层payload.body自然为空。
  • 解析出的多余符号多是HTML转义字符(如 、
)、冗余换行/制表符,或是HTML标签清理不彻底留下的残留内容。

解决方案

  1. 递归遍历邮件的parts结构,优先提取text/plain类型的正文,没有则把text/html转成纯文本;
  2. 用正则和字符串方法针对性清理多余符号和无效空白。

代码示例:

import re
from bs4 import BeautifulSoup
import base64
from googleapiclient.discovery import build

def get_email_body(payload):
    """递归提取邮件正文,优先取纯文本"""
    if 'parts' in payload:
        for part in payload['parts']:
            # 先找text/plain类型
            if part['mimeType'] == 'text/plain':
                data = part['body'].get('data', '')
                if data:
                    return base64.urlsafe_b64decode(data).decode('utf-8')
            # 没有纯文本就转HTML为文本
            elif part['mimeType'] == 'text/html':
                data = part['body'].get('data', '')
                if data:
                    html_content = base64.urlsafe_b64decode(data).decode('utf-8')
                    soup = BeautifulSoup(html_content, 'html.parser')
                    return soup.get_text(separator=' ')
            # 递归处理嵌套的parts
            nested_body = get_email_body(part)
            if nested_body:
                return nested_body
    # 处理单结构邮件
    if payload['mimeType'] == 'text/plain':
        data = payload['body'].get('data', '')
        return base64.urlsafe_b64decode(data).decode('utf-8') if data else ''
    elif payload['mimeType'] == 'text/html':
        data = payload['body'].get('data', '')
        if data:
            html_content = base64.urlsafe_b64decode(data).decode('utf-8')
            soup = BeautifulSoup(html_content, 'html.parser')
            return soup.get_text(separator=' ')
    return ''

def clean_body_text(text):
    """清理文本中的无效符号和冗余空白"""
    if not text:
        return ''
    # 移除HTML转义字符
    text = re.sub(r'&[a-zA-Z0-9#]+;', '', text)
    # 合并多余换行、制表符为单个空格
    text = re.sub(r'\s+', ' ', text).strip()
    return text

# 使用示例
service = build('gmail', 'v1', credentials=credentials)
message = service.users().messages().get(userId='me', id='目标邮件ID', format='full').execute()
raw_body = get_email_body(message['payload'])
cleaned_body = clean_body_text(raw_body)

问题2:添加解析优化代码后DataFrame为空

原因分析

  • 解析或清理函数返回了None而非空字符串,导致数据列表中无有效值;
  • 邮件遍历逻辑出错,没有正确把解析结果存入数据列表;
  • 清理后的正文全为空字符串,过滤后无数据残留。

解决方案

  1. 确保解析、清理函数始终返回字符串(即使无内容也返回'',禁止返回None);
  2. 调试时打印每一步的解析结果,确认有有效数据被收集;
  3. 构建DataFrame时明确列名,必要时保留空内容行再按需过滤。

代码示例:

import pandas as pd

# 收集邮件数据
emails_data = []
# 假设已获取邮件列表messages_list
for msg in messages_list:
    msg_id = msg['id']
    message = service.users().messages().get(userId='me', id=msg_id, format='full').execute()
    # 获取并清理正文
    raw_body = get_email_body(message['payload'])
    cleaned_body = clean_body_text(raw_body)
    # 提取邮件主题
    subject = next(h['value'] for h in message['payload']['headers'] if h['name'] == 'Subject')
    # 存入数据列表
    emails_data.append({
        '邮件ID': msg_id,
        '主题': subject,
        '正文': cleaned_body
    })

# 构建DataFrame
df = pd.DataFrame(emails_data)
# 打印前几行确认数据
print(df.head())
# 按需过滤空正文行
df = df[df['正文'] != ''].reset_index(drop=True)

内容的提问来源于stack exchange,提问作者UnKB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 01:20:23