You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python调用Gmail API批量获取指定邮件头并优化十万级邮件的标签处理

如何使用Python调用Gmail API批量获取指定邮件头并优化十万级邮件的标签处理

我完全懂你这种痛苦——15万+邮件单条处理每秒1-2个,要跑到天荒地老,而且Gmail API的批量文档确实太零散了,ChatGPT还总是给半吊子代码。我帮你把批量获取邮件头和批量修改标签的逻辑都理顺,直接给你能跑的完整方案。

一、先搞定批量获取指定邮件头的问题

你原来的batch代码主要问题是回调函数的参数不对,还有结果收集的逻辑没理清楚。Gmail API的batch请求需要正确的回调来捕获每一个请求的响应,而且要注意批量大小(官方建议最多100个请求一批,避免超时)。

正确的批量获取邮件头实现

这里我们用一个类来收集结果,比全局变量更稳妥,还能处理请求异常:

from __future__ import print_function
import os.path
import re
from google.auth.transport.requests import Request
from google.oauth2.credentials import Credentials
from google_auth_oauthlib.flow import InstalledAppFlow
from googleapiclient.discovery import build
from googleapiclient.errors import HttpError

SCOPES = ['https://www.googleapis.com/auth/gmail.modify',
          'https://www.googleapis.com/auth/gmail.labels']

# 用容器类收集批量请求结果,比全局变量更安全
class BatchResultContainer:
    def __init__(self):
        self.results = []
    
    def callback(self, request_id, response, exception):
        if exception:
            print(f"请求 {request_id} 失败: {str(exception)}")
        else:
            # 提取我们需要的头信息和邮件ID
            msg_id = response['id']
            headers = {h['name']: h['value'] for h in response['payload']['headers']}
            self.results.append({
                'id': msg_id,
                'From': headers.get('From', ''),
                'Reply-To': headers.get('Reply-To', ''),
                'Sender': headers.get('Sender', ''),
                'Return-Path': headers.get('Return-Path', '')
            })

def main():
    creds = None
    if os.path.exists('token.json'):
        creds = Credentials.from_authorized_user_file('token.json', SCOPES)
    if not creds or not creds.valid:
        if creds and creds.expired and creds.refresh_token:
            creds.refresh(Request())
        else:
            flow = InstalledAppFlow.from_client_secrets_file(
                'C:/path/to/credentials.json', SCOPES)
            creds = flow.run_local_server(port=0)
        with open('token.json', 'w') as token:
            token.write(creds.to_json())

    try:
        service = build('gmail', 'v1', credentials=creds)
        all_message_ids = []
        page_token = None

        # 分页拉取所有INBOX的邮件ID(maxResults=500是list接口的上限)
        while True:
            results = service.users().messages().list(
                userId='me', labelIds=['INBOX'], maxResults=500, pageToken=page_token).execute()
            messages = results.get('messages', [])
            all_message_ids.extend([msg['id'] for msg in messages])
            page_token = results.get('nextPageToken')
            if not page_token:
                break
        print(f"共获取到 {len(all_message_ids)} 封邮件ID")

        # 分批处理,每100个邮件一批(避免batch请求超时)
        container = BatchResultContainer()
        batch_size = 100
        for i in range(0, len(all_message_ids), batch_size):
            batch_ids = all_message_ids[i:i+batch_size]
            batch = service.new_batch_http_request(callback=container.callback)
            
            # 给每个邮件ID添加获取指定头的请求
            for msg_id in batch_ids:
                batch.add(service.users().messages().get(
                    userId='me', id=msg_id, format='metadata',
                    metadataHeaders=['From', 'Reply-To', 'Sender', 'Return-Path']
                ))
            
            print(f"正在处理第 {i//batch_size + 1} 批邮件,共 {len(batch_ids)} 封")
            batch.execute()
        
        # 这里container.results就是所有邮件的指定头信息了
        print(f"批量获取完成,共收集到 {len(container.results)} 封邮件的头信息")

        # 接下来就可以基于这些头信息处理标签了
        process_labels(service, container.results)

    except HttpError as error:
        print(f"API请求出错: {error}")

# 标签处理逻辑(下面会完善)
def process_labels(service, email_headers):
    pass

if __name__ == '__main__':
    main()

二、批量处理邮件标签(核心优化点)

获取到邮件头后,我们可以按域名分组,然后用批量修改标签的接口一次性处理同域名的所有邮件,这比单条modify快N倍。还要注意标签的层级创建(反序域名的层级标签)要复用已有的标签,避免重复创建。

完善标签处理函数

把你原来的标签创建逻辑和批量修改结合起来:

def get_sender_domain(email_headers):
    """从邮件头中提取发送者域名,并返回反序格式"""
    m = re.compile(r'^[^<]*<?([^@<> ]+@[^@<> ]+)>?$')
    sender_email = None
    # 按优先级提取发送者邮箱
    if email_headers['From'] and m.match(email_headers['From']):
        sender_email = m.match(email_headers['From']).group(1)
    elif email_headers['Reply-To'] and m.match(email_headers['Reply-To']):
        sender_email = m.match(email_headers['Reply-To']).group(1)
    elif email_headers['Return-Path'] and m.match(email_headers['Return-Path']):
        sender_email = m.match(email_headers['Return-Path']).group(1)
    elif email_headers['Sender'] and m.match(email_headers['Sender']):
        sender_email = m.match(email_headers['Sender']).group(1)
    
    if not sender_email:
        return 'NoSender'
    domain = sender_email.split('@')[-1]
    return '.'.join(reversed(domain.split('.')))

def create_label_path(rev_domain, service):
    """创建层级标签(如果不存在),返回最终标签ID"""
    domains = rev_domain.split('.')
    current_path = ''
    label_id = None
    # 先获取所有已存在的标签
    existing_labels = service.users().labels().list(userId='me').execute()
    label_map = {label['name']: label['id'] for label in existing_labels['labels']}

    for part in domains:
        current_path = f"{current_path}/{part}" if current_path else part
        if current_path not in label_map:
            # 创建新标签
            new_label = service.users().labels().create(
                userId='me', body={'name': current_path}).execute()
            label_map[current_path] = new_label['id']
        label_id = label_map[current_path]
    return label_id

def process_labels(service, email_headers):
    """按域名分组,批量修改邮件标签"""
    # 按反序域名分组邮件ID
    domain_email_map = {}
    for email in email_headers:
        rev_domain = get_sender_domain(email)
        if rev_domain not in domain_email_map:
            domain_email_map[rev_domain] = []
        domain_email_map[rev_domain].append(email['id'])
    
    # 批量处理每个域名的邮件
    for rev_domain, msg_ids in domain_email_map.items():
        print(f"处理域名 {rev_domain} 的 {len(msg_ids)} 封邮件")
        # 获取或创建标签ID
        label_id = create_label_path(rev_domain, service)
        # 批量修改标签(一次最多处理1000个邮件ID)
        batch_size = 1000
        for i in range(0, len(msg_ids), batch_size):
            batch_ids = msg_ids[i:i+batch_size]
            # 批量移除INBOX标签,添加目标标签
            service.users().messages().batchModify(
                userId='me',
                body={
                    'ids': batch_ids,
                    'removeLabelIds': ['INBOX'],
                    'addLabelIds': [label_id]
                }
            ).execute()
        print(f"完成域名 {rev_domain} 的标签修改")

三、关于Google Apps Script的补充说明

你提到的GAS确实有这个坑——它的GmailApp只能给线程加标签,而Gmail API的Python客户端支持给单个邮件加标签,所以Python方案是唯一能满足你需求的。

四、最终优化点

  1. 可以在批量处理时加入日志,方便追踪进度
  2. 如果邮件量极大,可以把邮件头信息临时写入本地文件,避免内存溢出
  3. 可以添加重试逻辑,处理偶尔的API请求失败

备注:内容来源于stack exchange,提问作者skeetastax

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 15:03:17