使用BeautifulSoup时出现UnicodeDecodeError('charmap')问题求助
解决BeautifulSoup读取HTML文件时的UnicodeDecodeError问题
问题根源
代码读取HTML文件时未指定编码,Windows系统默认用cp1252编码解析,但你的HTML文件包含utf-8编码的特殊字符(比如❤️),导致解码失败。重装BeautifulSoup或更换文件无法解决,因为问题出在文件读取的编码设置上。
修复方案
修改open()函数,明确指定编码为utf-8,匹配HTML文件的<meta charset="utf-8">声明:
from bs4 import BeautifulSoup with open("website.html", encoding='utf-8') as file: html_doc = file.read() soup = BeautifulSoup(html_doc, 'html.parser') print(soup.title.name)
临时替代方案(不推荐)
如果不确定文件编码,可添加errors='ignore'参数跳过无法解码的字符,但会丢失部分内容:
with open("website.html", errors='ignore') as file: html_doc = file.read()
内容的提问来源于stack exchange,提问作者Xareni Galindo
相关产品推荐
相关产品推荐

