Python获取UTF-8编码HTML标题乱码,求解决方法
解决HTML标题乱码问题
你的乱码问题根源是打开文件时未指定UTF-8编码——Python的open()函数默认使用系统编码读取文件,即便目标文档本身是UTF-8格式,也会因解码方式不匹配出现乱码。
直接修改open()函数,添加encoding='utf-8'参数即可解决:
from bs4 import BeautifulSoup with open("index.html", encoding='utf-8') as file: src = file.read() soup = BeautifulSoup(src, "lxml") title = soup.title.text print(title)
也可以通过读取字节流再手动解码的方式实现,效果一致:
from bs4 import BeautifulSoup with open("index.html", 'rb') as file: src = file.read().decode('utf-8') soup = BeautifulSoup(src, "lxml") title = soup.title.text print(title)
两种方法都能确保文件以UTF-8编码正确读取,从而获取正常的标题文本。
内容的提问来源于stack exchange,提问作者user18692348
相关产品推荐
相关产品推荐

