使用BeautifulSoup处理windows-1255编码希伯来语网页时编码出错求助
解决希伯来语网站编码导致的UnicodeEncodeError问题
嘿,我来帮你搞定这个编码问题!你遇到的UnicodeEncodeError主要是两个原因:一是requests可能没正确识别网站声明的windows-1255编码来解码内容,二是后续输出或保存时用了不支持希伯来语字符的默认编码(比如Windows系统的charmap)。
下面是直接可用的解决方案,替换你的代码就行:
from bs4 import BeautifulSoup import requests # 目标网站URL url = 'http://www.plonter.co.il' # 发送请求 r = requests.get(url) # 关键步骤:手动指定编码为网站声明的windows-1255 # 避免requests自动推断编码时出现偏差 r.encoding = 'windows-1255' # 用BeautifulSoup解析正确解码后的文本 soup = BeautifulSoup(r.text, 'html.parser') # 后续输出或保存时,统一用utf-8编码彻底规避报错 # 示例1:打印解析后的内容 print(soup.prettify(encoding='utf-8').decode('utf-8')) # 示例2:将内容保存为HTML文件 with open('plonter_content.html', 'w', encoding='utf-8') as file: file.write(soup.prettify())
为什么这么处理?
- 手动设置
r.encoding:虽然requests会自动尝试检测编码,但网站通过<meta>标签声明的windows-1255可能没被正确识别,手动指定能确保解码后的文本是准确的希伯来语内容。 - 用utf-8输出/保存:
utf-8支持所有Unicode字符,而Windows默认的charmap编码不包含部分希伯来语特殊字符,用utf-8可以彻底避免character maps to <undefined>的错误。
如果还是遇到问题,可以尝试手动解码响应二进制内容再传给BeautifulSoup:
# 替代r.text的方式:手动解码二进制响应 html_content = r.content.decode('windows-1255') soup = BeautifulSoup(html_content, 'html.parser')
内容的提问来源于stack exchange,提问作者Idan
相关产品推荐
相关产品推荐

