如何使用Gmail API(Python)获取完整邮件正文
使用Gmail API提取完整邮件正文的Python实现
没问题!我来帮你搞定Gmail API提取完整邮件正文的事儿。你说得对,snippet字段确实只返回简短预览,要拿到完整内容得从邮件的payload结构里解析——Gmail的邮件 payload 分纯文本和HTML两种格式,部分邮件还会是多部分(multipart)结构,得针对性处理。
咱先说说前置准备:
- 先在Google Cloud Console里启用Gmail API,下载
credentials.json文件到项目目录 - 安装依赖包:
pip install google-api-python-client google-auth-httplib2 google-auth-oauthlib
接下来是完整的Python示例代码,包含授权、邮件列表获取和正文解析:
import base64 from googleapiclient.discovery import build from google_auth_oauthlib.flow import InstalledAppFlow from google.auth.transport.requests import Request import pickle import os.path # 授权范围,这里用只读权限足够 SCOPES = ['https://www.googleapis.com/auth/gmail.readonly'] def get_gmail_service(): creds = None # 检查是否有保存的token if os.path.exists('token.pickle'): with open('token.pickle', 'rb') as token: creds = pickle.load(token) # 如果没有有效凭据,重新登录 if not creds or not creds.valid: if creds and creds.expired and creds.refresh_token: creds.refresh(Request()) else: flow = InstalledAppFlow.from_client_secrets_file( 'credentials.json', SCOPES) creds = flow.run_local_server(port=0) # 保存token供下次使用 with open('token.pickle', 'wb') as token: pickle.dump(creds, token) # 构建Gmail服务对象 return build('gmail', 'v1', credentials=creds) def get_email_body(message): """解析邮件的完整正文""" body = "" payload = message['payload'] # 处理多部分邮件(比如同时有纯文本和HTML版本) if 'parts' in payload: # 遍历所有parts,优先取纯文本,也可以根据需求取HTML for part in payload['parts']: if part['mimeType'] == 'text/plain': body = base64.urlsafe_b64decode(part['body']['data']).decode('utf-8') break elif part['mimeType'] == 'text/html': # 如果需要HTML格式,就用这个分支 # body = base64.urlsafe_b64decode(part['body']['data']).decode('utf-8') pass # 处理单部分邮件 else: if payload['mimeType'] == 'text/plain' or payload['mimeType'] == 'text/html': body = base64.urlsafe_b64decode(payload['body']['data']).decode('utf-8') return body def main(): service = get_gmail_service() # 获取最近10封邮件的ID(可以修改maxResults参数调整数量) results = service.users().messages().list(userId='me', maxResults=10).execute() messages = results.get('messages', []) if not messages: print('No messages found.') else: print('Messages:') for message in messages: # 获取完整邮件详情 msg = service.users().messages().get(userId='me', id=message['id']).execute() # 解析正文 body = get_email_body(msg) # 获取邮件主题 subject = [header['value'] for header in msg['payload']['headers'] if header['name'] == 'Subject'][0] print(f"\n--- 主题: {subject} ---") print(body) if __name__ == '__main__': main()
关键逻辑说明:
- 授权部分:用Google的OAuth2流程,首次运行会弹出浏览器让你登录授权,之后会保存
token.pickle,下次直接用 - 正文解析:
- 先判断邮件是
multipart(带parts字段)还是单部分 - 多部分邮件里,我们优先提取
text/plain格式的正文,如果你需要HTML格式,注释掉纯文本的break,启用HTML分支即可 - Gmail返回的正文是base64url编码的,所以要用
base64.urlsafe_b64decode解码,再转成UTF-8字符串
- 先判断邮件是
- 获取邮件主题:从
payload['headers']里筛选name为Subject的字段
你可以根据自己的需求调整,比如修改users().messages().list()的参数来过滤特定邮件(比如按标签、日期),或者调整正文的解析逻辑优先取HTML。
内容的提问来源于stack exchange,提问作者Shreya Sharma
相关产品推荐
相关产品推荐

