Gmail邮件Base64解码异常求助(附修正后Python代码)
问题:Gmail API邮件正文解码异常处理
更新说明
感谢Tuppitappi,其在相关问题中提到的urlsafe_b64decode正是所需方法!
问题描述
我编写了一个Python脚本,用于拉取Gmail邮件、解析为类列表并打印结果。除Base64解码环节外,其余功能均正常:部分输出完全无法打印,部分转换后HTML正文中的元素仍处于编码状态。我尝试了多种编码方式,包括从邮件头中获取的编码,但每次都报错,提示输入字符串无效或格式不正确,请问该如何正确解码邮件正文?
相关解码代码
我从邮件部分获取编码方式,解码逻辑如下:
try: newOne.body = base64.urlsafe_b64decode(newOne.body).decode(newOne.charSet, 'backslashreplace') except Exception as defaultError: print("Error with default decoding: ", defaultError) try: newOne.body = base64.urlsafe_b64decode(newOne.body).decode("iso8859_2", 'backslashreplace') except Exception as isoError: print("Error with ISO decoding: ", isoError) try: # UTF-8 is the default newOne.body = base64.urlsafe_b64decode(newOne.body).decode('utf-8', 'backslashreplace') except Exception as utfError: print("Error with UTF decoding: ", utfError) try: newOne.body = base64.urlsafe_b64decode(newOne.body).decode("ascii", 'backslashreplace') except Exception as asciiError: print("Error with ASCII decoding: ", asciiError) try: newOne.body = base64.urlsafe_b64decode(newOne.body).decode("cp437", 'backslashreplace') except Exception as oldSchoolError: print("Error with Old School decoding: ", oldSchoolError) newOne.body = newOne.body
错误示例
运行时出现的错误信息如下:
Error with default decoding: 'utf-8' codec can't decode byte 0xdc in position 184: invalid continuation byte Error with default decoding: Invalid base64-encoded string: number of data characters (4701) cannot be 1 more than a multiple of 4 Error with ISO decoding: Invalid base64-encoded string: number of data characters (4701) cannot be 1 more than a multiple of 4 Error with ASCII decoding: Invalid base64-encoded string: number of data characters (4701) cannot be 1 more than a multiple of 4 Error with UTF decoding: Invalid base64-encoded string: number of data characters (4701) cannot be 1 more than a multiple of 4 Error with default decoding: Incorrect padding Error with ISO decoding: Incorrect padding Error with ASCII decoding: Incorrect padding Error with UTF decoding: Incorrect padding Error with default decoding: 'utf-8' codec can't decode byte 0xa0 in position 20: invalid start byte
已修正的完整代码
import binascii import os.path from datetime import datetime import base64 import dateutil.parser from google.auth.transport.requests import Request from google.oauth2.credentials import Credentials from google_auth_oauthlib.flow import InstalledAppFlow from googleapiclient.discovery import build from googleapiclient.errors import HttpError from typing import List # If modifying these scopes, delete the file token.json. SCOPES = ["https://www.googleapis.com/auth/gmail.readonly"] class Attachment(): attachmentId: str mimeType: str fileName: str class NewMail(): mailId: str subject: str sentTo: str sentFrom: str dateSent: datetime body: str contentType: str charSet: str attachments: List[Attachment] def main(): """Shows basic usage of the Gmail API. Lists the user's Gmail labels. """ creds = None # The file token.json stores the user's access and refresh tokens, and is # created automatically when the authorization flow completes for the first # time. if os.path.exists("token.json"): creds = Credentials.from_authorized_user_file("token.json", SCOPES) # If there are no (valid) credentials available, let the user log in. if not creds or not creds.valid: if creds and creds.expired and creds.refresh_token: creds.refresh(Request()) else: flow = InstalledAppFlow.from_client_secrets_file( "credentials.json", SCOPES ) creds = flow.run_local_server(port=56560) # Save the credentials for the next run with open("token.json", "w") as token: token.write(creds.to_json()) # Call the Gmail API service = build("gmail", "v1", credentials=creds) try: results = service.users().messages().list(userId="me",maxResults=10).execute() messages = results.get("messages", []) if not messages: print("No message found.") return currentMail = list() print("Messages:") for message in messages: newAttachments = list() newOne = NewMail() newOne.mailId = message.get("id") newOne.body = "" newOne.charSet = "utf_8" thsMsg = service.users().messages().get(userId="me",id=message.get("id")).execute() #First we process the Header for the main email for header in thsMsg.get("payload").get("headers"): if header["name"] == "Subject": newOne.subject = header["value"] elif header["name"] == "To": newOne.sentTo = header["value"] elif header["name"] == "From": newOne.sentFrom = header["value"] elif header["name"] == "To": newOne.dateSent = dateutil.parser.parse(header["value"]) elif header["name"] == "Content-Type": # We need to extract out the parts we want typeParts = header["value"].split(";") for typePart in typeParts: if "charset" in typePart: newOne.charSet = typePart.replace("charset=","") elif "/" in typePart: newOne.contentType = typePart #Next we look to see if the body has anything in it if thsMsg.get("payload").get("body") is not None and thsMsg.get("payload").get("body").get("data") is not None: newOne.body = thsMsg.get("payload").get("body").get("data").strip() #Finally we process the multipart - looking for attachments if thsMsg.get("payload").get("parts") is not None: for attachMe in thsMsg.get("payload").get("parts"): if attachMe.get("filename") is not None and attachMe.get("filename") != "": attachThis = Attachment() attachThis.attachmentId = attachMe.get("partId"), attachThis.mimeType = attachMe.get("mimType"), attachThis.fileName = attachMe.get("filename"), newOne.attachments.append(attachThis) elif (newOne.body == "" and attachMe.get("body").get("data") is not None and attachMe.get("body").get("data") != ""): newOne.body = attachMe.get("body").get("data").strip() #We also grab this version encoding and content type for header in attachMe.get("headers"): if header["name"] == "Content-Type": # We need to extract out the parts we want typeParts = header["value"].split(";") for typePart in typeParts: if "charset" in typePart: newOne.charSet = typePart.replace("charset=", "") elif "/" in typePart: newOne.contentType = typePart break #GMAIL is base64 encoded so we need to decode it if newOne.body != "": try: try: newOne.body = base64.urlsafe_b64decode(newOne.body).decode(newOne.charSet, 'backslashreplace') except Exception as defaultError: print("Error with default decoding: ", defaultError) try: newOne.body = base64.urlsafe_b64decode(newOne.body).decode("iso8859_2", 'backslashreplace') except Exception as isoError: print("Error with ISO decoding: ", isoError) try: # UTF-8 is the default newOne.body = base64.urlsafe_b64decode(newOne.body).decode('utf-8', 'backslashreplace') except Exception as utfError: print("Error with UTF decoding: ", utfError) try: newOne.body = base64.urlsafe_b64decode(newOne.body).decode("ascii", 'backslashreplace') except Exception as asciiError: print("Error with ASCII decoding: ", asciiError) try: newOne.body = base64.urlsafe_b64decode(newOne.body).decode("cp437", 'backslashreplace') except Exception as oldSchoolError: print("Error with Old School decoding: ", oldSchoolError) newOne.body = newOne.body currentMail.append(newOne) #Test print out to verify email content for thisMail in currentMail: print("To: " + thisMail.sentTo + "\n") print("From: " + thisMail.sentFrom + "\n") print("Subject: " + thisMail.subject + "\n") if thisMail.body != "": print("Body: " + thisMail.body) else : print("No email body") print("---------------------------------------\n") except HttpError as error: # TODO(developer) - Handle errors from gmail API. print(f"An error occurred: {error}") finally: service.close() if __name__ == "__main__": main()
尝试过的方法
我试过多种编码方式,包括查阅Python标准编码文档后测试的不同选项,也直接使用邮件头里的编码,但都无效,每次都报输入无效或格式错误。
解决方法
- 处理Base64填充问题:Gmail返回的base64数据可能缺少必要的填充字符
=,解码前先补全:
# 补全base64填充 padding = len(newOne.body) % 4 if padding != 0: newOne.body += '=' * (4 - padding)
将这段代码放在urlsafe_b64decode之前,解决"Incorrect padding"和数据长度不符的问题。
- 修复字符集提取逻辑:当前提取charset时未处理空格,需去掉多余空格:
if "charset" in typePart: newOne.charSet = typePart.replace("charset=", "").strip()
- 处理HTML实体编码:解码后若仍有HTML元素编码(如
&),用html.unescape还原:
import html # 解码后处理HTML实体 newOne.body = html.unescape(newOne.body)
- 简化解码逻辑:无需多层嵌套try-except,先统一处理base64问题,再用指定charset解码,失败后用
errors='replace'兜底:
if newOne.body != "": # 处理填充 padding = len(newOne.body) % 4 if padding != 0: newOne.body += '=' * (4 - padding) try: decoded_bytes = base64.urlsafe_b64decode(newOne.body) # 先尝试指定charset newOne.body = decoded_bytes.decode(newOne.charSet.strip(), 'replace') except: # 兜底用utf-8 newOne.body = decoded_bytes.decode('utf-8', 'replace') # 处理HTML实体 newOne.body = html.unescape(newOne.body)
- 修正日期解析BUG:当前代码重复判断
"To"头,应改为"Date"头:
elif header["name"] == "Date": newOne.dateSent = dateutil.parser.parse(header["value"])
- 修复附件MIME类型拼写错误:
attachMe.get("mimType")应改为attachMe.get("mimeType")。
内容的提问来源于stack exchange,提问作者user23640624
相关产品推荐
相关产品推荐

