You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Gmail API读取邮件时的编码解码及HTML标签异常问题

解决Gmail API读取邮件时的编码与HTML标签问题

嘿,这两个问题我之前用Gmail API的时候也踩过坑,给你分享下靠谱的解决办法:

一、匈牙利语重音字符显示异常(比如G=C3=A1bor变成Gábor)

这是因为Gmail API返回的内容用了Quoted-Printable编码,这种编码会把非ASCII字符转成=C3=A1这类格式,你需要对内容做解码处理:

方法1:用Python内置quopri模块直接解码

如果只是处理单个字符串片段,可以直接用这个模块:

import quopri

encoded_str = "G=C3=A1bor"
decoded_str = quopri.decodestring(encoded_str).decode('utf-8')
print(decoded_str)  # 输出: Gábor

方法2:用email库解析完整邮件(更推荐)

Gmail返回的邮件结构通常是多部分的,用email库能自动处理编码逻辑,避免手动踩坑:

from email import message_from_bytes
import base64

# 假设你已经通过API拿到了原始邮件内容
raw_message = service.users().messages().get(userId='me', id=message_id, format='raw').execute()
msg_bytes = base64.urlsafe_b64decode(raw_message['raw'])

# 解析邮件对象
msg = message_from_bytes(msg_bytes)

# 遍历邮件各部分,获取正确编码的内容
for part in msg.walk():
    if part.get_content_type() == 'text/plain':
        plain_text = part.get_payload(decode=True).decode(part.get_content_charset() or 'utf-8')
        print(plain_text)
    elif part.get_content_type() == 'text/html':
        html_content = part.get_payload(decode=True).decode(part.get_content_charset() or 'utf-8')
        print(html_content)

二、HTML标签损坏问题

HTML标签损坏大多是因为直接读取了未解码的原始内容,或者手动处理字符串时破坏了结构。用上面的email库解析HTML部分是最稳妥的方式,要是还需要修复或格式化HTML,可以搭配BeautifulSoup:

from bs4 import BeautifulSoup

# 接上面的代码,拿到html_content后
soup = BeautifulSoup(html_content, 'html.parser')
# 自动修复损坏标签并格式化
formatted_html = soup.prettify()
print(formatted_html)

这样处理后,重音字符能正常显示,HTML标签也会保持完整结构。

内容的提问来源于stack exchange,提问作者Bar6

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:11:20