You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup时出现UnicodeDecodeError('charmap')问题求助

解决BeautifulSoup读取HTML文件时的UnicodeDecodeError问题

问题根源

代码读取HTML文件时未指定编码,Windows系统默认用cp1252编码解析,但你的HTML文件包含utf-8编码的特殊字符(比如❤️),导致解码失败。重装BeautifulSoup或更换文件无法解决,因为问题出在文件读取的编码设置上。

修复方案

修改open()函数,明确指定编码为utf-8,匹配HTML文件的<meta charset="utf-8">声明:

from bs4 import BeautifulSoup

with open("website.html", encoding='utf-8') as file:
    html_doc = file.read()

soup = BeautifulSoup(html_doc, 'html.parser')
print(soup.title.name)

临时替代方案(不推荐)

如果不确定文件编码,可添加errors='ignore'参数跳过无法解码的字符,但会丢失部分内容:

with open("website.html", errors='ignore') as file:
    html_doc = file.read()

内容的提问来源于stack exchange,提问作者Xareni Galindo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 14:27:16