使用Gmail API读取邮件时的编码解码及HTML标签异常问题
解决Gmail API读取邮件时的编码与HTML标签问题
嘿,这两个问题我之前用Gmail API的时候也踩过坑,给你分享下靠谱的解决办法:
一、匈牙利语重音字符显示异常(比如G=C3=A1bor变成Gábor)
这是因为Gmail API返回的内容用了Quoted-Printable编码,这种编码会把非ASCII字符转成=C3=A1这类格式,你需要对内容做解码处理:
方法1:用Python内置quopri模块直接解码
如果只是处理单个字符串片段,可以直接用这个模块:
import quopri encoded_str = "G=C3=A1bor" decoded_str = quopri.decodestring(encoded_str).decode('utf-8') print(decoded_str) # 输出: Gábor
方法2:用email库解析完整邮件(更推荐)
Gmail返回的邮件结构通常是多部分的,用email库能自动处理编码逻辑,避免手动踩坑:
from email import message_from_bytes import base64 # 假设你已经通过API拿到了原始邮件内容 raw_message = service.users().messages().get(userId='me', id=message_id, format='raw').execute() msg_bytes = base64.urlsafe_b64decode(raw_message['raw']) # 解析邮件对象 msg = message_from_bytes(msg_bytes) # 遍历邮件各部分,获取正确编码的内容 for part in msg.walk(): if part.get_content_type() == 'text/plain': plain_text = part.get_payload(decode=True).decode(part.get_content_charset() or 'utf-8') print(plain_text) elif part.get_content_type() == 'text/html': html_content = part.get_payload(decode=True).decode(part.get_content_charset() or 'utf-8') print(html_content)
二、HTML标签损坏问题
HTML标签损坏大多是因为直接读取了未解码的原始内容,或者手动处理字符串时破坏了结构。用上面的email库解析HTML部分是最稳妥的方式,要是还需要修复或格式化HTML,可以搭配BeautifulSoup:
from bs4 import BeautifulSoup # 接上面的代码,拿到html_content后 soup = BeautifulSoup(html_content, 'html.parser') # 自动修复损坏标签并格式化 formatted_html = soup.prettify() print(formatted_html)
这样处理后,重音字符能正常显示,HTML标签也会保持完整结构。
内容的提问来源于stack exchange,提问作者Bar6
相关产品推荐
相关产品推荐

