You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup处理windows-1255编码希伯来语网页时编码出错求助

解决希伯来语网站编码导致的UnicodeEncodeError问题

嘿,我来帮你搞定这个编码问题!你遇到的UnicodeEncodeError主要是两个原因:一是requests可能没正确识别网站声明的windows-1255编码来解码内容,二是后续输出或保存时用了不支持希伯来语字符的默认编码(比如Windows系统的charmap)。

下面是直接可用的解决方案,替换你的代码就行:

from bs4 import BeautifulSoup
import requests

# 目标网站URL
url = 'http://www.plonter.co.il'
# 发送请求
r = requests.get(url)

# 关键步骤:手动指定编码为网站声明的windows-1255
# 避免requests自动推断编码时出现偏差
r.encoding = 'windows-1255'

# 用BeautifulSoup解析正确解码后的文本
soup = BeautifulSoup(r.text, 'html.parser')

# 后续输出或保存时,统一用utf-8编码彻底规避报错
# 示例1:打印解析后的内容
print(soup.prettify(encoding='utf-8').decode('utf-8'))

# 示例2:将内容保存为HTML文件
with open('plonter_content.html', 'w', encoding='utf-8') as file:
    file.write(soup.prettify())

为什么这么处理?

  • 手动设置r.encoding:虽然requests会自动尝试检测编码,但网站通过<meta>标签声明的windows-1255可能没被正确识别,手动指定能确保解码后的文本是准确的希伯来语内容。
  • 用utf-8输出/保存:utf-8支持所有Unicode字符,而Windows默认的charmap编码不包含部分希伯来语特殊字符,用utf-8可以彻底避免character maps to <undefined>的错误。

如果还是遇到问题,可以尝试手动解码响应二进制内容再传给BeautifulSoup:

# 替代r.text的方式:手动解码二进制响应
html_content = r.content.decode('windows-1255')
soup = BeautifulSoup(html_content, 'html.parser')

内容的提问来源于stack exchange,提问作者Idan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:41:33