使用Python调用Gmail API无法获取完整邮件正文求助
解决Gmail API无法获取完整邮件正文的问题
你的问题我看明白了——现在代码里拿的是邮件的snippet字段,这玩意儿本来就是Gmail自动生成的邮件摘要,天生就是截断的,不管你请求时用Raw还是Full格式,它都只会返回部分内容。要拿到完整正文,得从邮件的payload结构里去提取,还要处理Base64编码的内容,因为Gmail API返回的邮件内容大多是经过Base64 URL安全编码的。
先看你提供的原始代码核心问题:
# 这里打印的是自动截断的摘要,不是完整正文 print(msg['snippet'])
下面是修改后的完整代码,我帮你调整了邮件正文的提取逻辑,还加了主题过滤、纯文本优先解析的处理,甚至帮你加了验证码提取的示例:
from __future__ import print_function import pickle import os.path import re from googleapiclient.discovery import build from google_auth_oauthlib.flow import InstalledAppFlow from google.auth.transport.requests import Request from bs4 import BeautifulSoup import base64 SCOPES = ['https://www.googleapis.com/auth/gmail.readonly'] def main(): """Shows basic usage of the Gmail API. Fetches full message content and extracts verification code. """ creds = None if os.path.exists('token.pickle'): with open('token.pickle', 'rb') as token: creds = pickle.load(token) # Handle credential refresh/login if not creds or not creds.valid: if creds and creds.expired and creds.refresh_token: creds.refresh(Request()) else: flow = InstalledAppFlow.from_client_secrets_file( 'credentials.json', SCOPES) creds = flow.run_local_server(port=0) with open('token.pickle', 'wb') as token: pickle.dump(creds, token) service = build('gmail', 'v1', credentials=creds) # 过滤主题含“验证码”的邮件,减少不必要的遍历 results = service.users().messages().list( userId='me', labelIds=['INBOX'], q='subject:验证码' ).execute() messages = results.get('messages', []) if not messages: print("No verification code messages found.") else: print("Processing verification messages:\n") for message in messages: # 指定format='full'确保获取完整邮件结构 msg = service.users().messages().get( userId='me', id=message['id'], format='full' ).execute() payload = msg['payload'] headers = payload['headers'] # 获取邮件主题,方便识别目标邮件 subject = "Unidentified Subject" for header in headers: if header['name'] == 'Subject': subject = header['value'] break print(f"=== {subject} ===") full_body = "" # 处理多部分邮件(绝大多数邮件的结构) if 'parts' in payload: for part in payload['parts']: # 优先提取纯文本格式,避免HTML标签干扰 if part['mimeType'] == 'text/plain': # 转换Base64 URL安全编码为标准编码 body_data = part['body']['data'].replace('-', '+').replace('_', '/') full_body = base64.b64decode(body_data).decode('utf-8') break # 没有纯文本时,解析HTML转成纯文本 elif part['mimeType'] == 'text/html': html_data = part['body']['data'].replace('-', '+').replace('_', '/') html_content = base64.b64decode(html_data).decode('utf-8') soup = BeautifulSoup(html_content, 'html.parser') full_body = soup.get_text(strip=False) # 处理单部分邮件(少数情况) else: body_data = payload['body']['data'].replace('-', '+').replace('_', '/') full_body = base64.b64decode(body_data).decode('utf-8') # 打印完整正文 print(full_body) # 提取一次性验证码(根据你的邮件内容定制的正则) code_match = re.search(r'您的一次性验证码:(\d+)', full_body) if code_match: print(f"\nExtracted verification code: {code_match.group(1)}\n") if __name__ == '__main__': main()
几个关键细节说明:
- 我加了
q='subject:验证码'过滤条件,直接定位目标邮件,不用遍历整个收件箱 - 优先提取
text/plain格式的正文,拿到的内容和你预期的完全一致,没有HTML标签干扰 - 处理了Base64 URL安全编码的转换:Gmail会把标准Base64的
+换成-、/换成_,必须替换回来才能正确解码 - 最后加了正则提取验证码的逻辑,直接帮你拿到需要的核心内容
内容的提问来源于stack exchange,提问作者ommse2etest Govid
相关产品推荐
相关产品推荐

