You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gmail邮件Base64解码异常求助(附修正后Python代码)

问题:Gmail API邮件正文解码异常处理

更新说明

感谢Tuppitappi,其在相关问题中提到的urlsafe_b64decode正是所需方法!

问题描述

我编写了一个Python脚本,用于拉取Gmail邮件、解析为类列表并打印结果。除Base64解码环节外,其余功能均正常:部分输出完全无法打印,部分转换后HTML正文中的元素仍处于编码状态。我尝试了多种编码方式,包括从邮件头中获取的编码,但每次都报错,提示输入字符串无效或格式不正确,请问该如何正确解码邮件正文?

相关解码代码

我从邮件部分获取编码方式,解码逻辑如下:

try:
    newOne.body = base64.urlsafe_b64decode(newOne.body).decode(newOne.charSet, 'backslashreplace')
except Exception as defaultError:
    print("Error with default decoding: ", defaultError)
    try:
        newOne.body = base64.urlsafe_b64decode(newOne.body).decode("iso8859_2", 'backslashreplace')
    except Exception as isoError:
        print("Error with ISO decoding: ", isoError)
        try:
            # UTF-8 is the default
            newOne.body = base64.urlsafe_b64decode(newOne.body).decode('utf-8', 'backslashreplace')
        except Exception as utfError:
            print("Error with UTF decoding: ", utfError)
            try:
                newOne.body = base64.urlsafe_b64decode(newOne.body).decode("ascii", 'backslashreplace')
            except Exception as asciiError:
                print("Error with ASCII decoding: ", asciiError)
                try:
                    newOne.body = base64.urlsafe_b64decode(newOne.body).decode("cp437", 'backslashreplace')
                except Exception as oldSchoolError:
                    print("Error with Old School decoding: ", oldSchoolError)
                    newOne.body = newOne.body

错误示例

运行时出现的错误信息如下:

Error with default decoding:  'utf-8' codec can't decode byte 0xdc in position 184: invalid continuation byte
Error with default decoding:  Invalid base64-encoded string: number of data characters (4701) cannot be 1 more than a multiple of 4
Error with ISO decoding:  Invalid base64-encoded string: number of data characters (4701) cannot be 1 more than a multiple of 4
Error with ASCII decoding:  Invalid base64-encoded string: number of data characters (4701) cannot be 1 more than a multiple of 4
Error with UTF decoding:  Invalid base64-encoded string: number of data characters (4701) cannot be 1 more than a multiple of 4
Error with default decoding:  Incorrect padding
Error with ISO decoding:  Incorrect padding
Error with ASCII decoding:  Incorrect padding
Error with UTF decoding:  Incorrect padding
Error with default decoding:  'utf-8' codec can't decode byte 0xa0 in position 20: invalid start byte

已修正的完整代码

import binascii
import os.path
from datetime import datetime
import base64
import dateutil.parser
from google.auth.transport.requests import Request
from google.oauth2.credentials import Credentials
from google_auth_oauthlib.flow import InstalledAppFlow
from googleapiclient.discovery import build
from googleapiclient.errors import HttpError
from typing import List

# If modifying these scopes, delete the file token.json.
SCOPES = ["https://www.googleapis.com/auth/gmail.readonly"]

class Attachment():
  attachmentId: str
  mimeType: str
  fileName: str
class NewMail():
  mailId: str
  subject: str
  sentTo: str
  sentFrom: str
  dateSent: datetime
  body: str
  contentType: str
  charSet: str
  attachments: List[Attachment]
def main():
  """Shows basic usage of the Gmail API.
  Lists the user's Gmail labels.
  """
  creds = None
  # The file token.json stores the user's access and refresh tokens, and is
  # created automatically when the authorization flow completes for the first
  # time.
  if os.path.exists("token.json"):
    creds = Credentials.from_authorized_user_file("token.json", SCOPES)
  # If there are no (valid) credentials available, let the user log in.
  if not creds or not creds.valid:
    if creds and creds.expired and creds.refresh_token:
      creds.refresh(Request())
    else:
      flow = InstalledAppFlow.from_client_secrets_file(
          "credentials.json", SCOPES
      )
      creds = flow.run_local_server(port=56560)
    # Save the credentials for the next run
    with open("token.json", "w") as token:
      token.write(creds.to_json())

  # Call the Gmail API
  service = build("gmail", "v1", credentials=creds)
  try:
    results = service.users().messages().list(userId="me",maxResults=10).execute()
    messages = results.get("messages", [])

    if not messages:
      print("No message found.")
      return

    currentMail = list()

    print("Messages:")
    for message in messages:
      newAttachments = list()
      newOne = NewMail()
      newOne.mailId = message.get("id")
      newOne.body = ""
      newOne.charSet = "utf_8"
      thsMsg = service.users().messages().get(userId="me",id=message.get("id")).execute()

      #First we process the Header for the main email
      for header in thsMsg.get("payload").get("headers"):
        if header["name"] == "Subject":
          newOne.subject = header["value"]
        elif header["name"] == "To":
          newOne.sentTo = header["value"]
        elif header["name"] == "From":
          newOne.sentFrom = header["value"]
        elif header["name"] == "To":
          newOne.dateSent = dateutil.parser.parse(header["value"])
        elif header["name"] == "Content-Type":
          # We need to extract out the parts we want
          typeParts = header["value"].split(";")
          for typePart in typeParts:
            if "charset" in typePart:
              newOne.charSet = typePart.replace("charset=","")
            elif "/" in typePart:
              newOne.contentType = typePart

      #Next we look to see if the body has anything in it
      if thsMsg.get("payload").get("body") is not None and thsMsg.get("payload").get("body").get("data") is not None:
        newOne.body = thsMsg.get("payload").get("body").get("data").strip()

      #Finally we process the multipart - looking for attachments
      if thsMsg.get("payload").get("parts") is not None:
        for attachMe in thsMsg.get("payload").get("parts"):
          if attachMe.get("filename") is not None and attachMe.get("filename") != "":
            attachThis = Attachment()
            attachThis.attachmentId = attachMe.get("partId"),
            attachThis.mimeType = attachMe.get("mimType"),
            attachThis.fileName = attachMe.get("filename"),
            newOne.attachments.append(attachThis)
          elif (newOne.body == "" and attachMe.get("body").get("data") is not None and
                attachMe.get("body").get("data") != ""):
            newOne.body = attachMe.get("body").get("data").strip()
            #We also grab this version encoding and content type
            for header in attachMe.get("headers"):
              if header["name"] == "Content-Type":
                # We need to extract out the parts we want
                typeParts = header["value"].split(";")
                for typePart in typeParts:
                  if "charset" in typePart:
                    newOne.charSet = typePart.replace("charset=", "")
                  elif "/" in typePart:
                    newOne.contentType = typePart
                break

      #GMAIL is base64 encoded so we need to decode it
      if newOne.body != "":
        try:
          try:
            newOne.body = base64.urlsafe_b64decode(newOne.body).decode(newOne.charSet, 'backslashreplace')
          except Exception as defaultError:
              print("Error with default decoding: ", defaultError)
              try:
                  newOne.body = base64.urlsafe_b64decode(newOne.body).decode("iso8859_2", 'backslashreplace')
              except Exception as isoError:
                  print("Error with ISO decoding: ", isoError)
                  try:
                      # UTF-8 is the default
                      newOne.body = base64.urlsafe_b64decode(newOne.body).decode('utf-8', 'backslashreplace')
                  except Exception as utfError:
                      print("Error with UTF decoding: ", utfError)
                      try:
                          newOne.body = base64.urlsafe_b64decode(newOne.body).decode("ascii", 'backslashreplace')
                      except Exception as asciiError:
                          print("Error with ASCII decoding: ", asciiError)
                          try:
                              newOne.body = base64.urlsafe_b64decode(newOne.body).decode("cp437", 'backslashreplace')
                          except Exception as oldSchoolError:
                              print("Error with Old School decoding: ", oldSchoolError)
                              newOne.body = newOne.body
      currentMail.append(newOne)

    #Test print out to verify email content
    for thisMail in currentMail:
      print("To: " + thisMail.sentTo + "\n")
      print("From: " + thisMail.sentFrom + "\n")
      print("Subject: " + thisMail.subject + "\n")
      if thisMail.body != "":
        print("Body: " + thisMail.body)
      else :
          print("No email body")
      print("---------------------------------------\n")

  except HttpError as error:
    # TODO(developer) - Handle errors from gmail API.
    print(f"An error occurred: {error}")
  finally:
    service.close()

if __name__ == "__main__":
  main()

尝试过的方法

我试过多种编码方式,包括查阅Python标准编码文档后测试的不同选项,也直接使用邮件头里的编码,但都无效,每次都报输入无效或格式错误。


解决方法

  1. 处理Base64填充问题:Gmail返回的base64数据可能缺少必要的填充字符=,解码前先补全:
# 补全base64填充
padding = len(newOne.body) % 4
if padding != 0:
    newOne.body += '=' * (4 - padding)

将这段代码放在urlsafe_b64decode之前,解决"Incorrect padding"和数据长度不符的问题。

  1. 修复字符集提取逻辑:当前提取charset时未处理空格,需去掉多余空格:
if "charset" in typePart:
    newOne.charSet = typePart.replace("charset=", "").strip()
  1. 处理HTML实体编码:解码后若仍有HTML元素编码(如&),用html.unescape还原:
import html
# 解码后处理HTML实体
newOne.body = html.unescape(newOne.body)
  1. 简化解码逻辑:无需多层嵌套try-except,先统一处理base64问题,再用指定charset解码,失败后用errors='replace'兜底:
if newOne.body != "":
    # 处理填充
    padding = len(newOne.body) % 4
    if padding != 0:
        newOne.body += '=' * (4 - padding)
    try:
        decoded_bytes = base64.urlsafe_b64decode(newOne.body)
        # 先尝试指定charset
        newOne.body = decoded_bytes.decode(newOne.charSet.strip(), 'replace')
    except:
        # 兜底用utf-8
        newOne.body = decoded_bytes.decode('utf-8', 'replace')
    # 处理HTML实体
    newOne.body = html.unescape(newOne.body)
  1. 修正日期解析BUG:当前代码重复判断"To"头,应改为"Date"头:
elif header["name"] == "Date":
    newOne.dateSent = dateutil.parser.parse(header["value"])
  1. 修复附件MIME类型拼写错误:attachMe.get("mimType")应改为attachMe.get("mimeType")。

内容的提问来源于stack exchange,提问作者user23640624

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 22:47:05