使用requests爬虫UTF-8编码下日语字符无法正常显示的问题求助
问题原因
- requests默认对该站点的响应编码识别错误,识别结果为不支持东亚字符的
ISO-8859-1,即使直接传递字节流给BeautifulSoup,库本身的自动编码检测也可能出现偏差,最终导致日语、韩文字符解析为乱码。 - 如果你使用的Windows系统默认终端编码为GBK,也会在打印输出环节出现二次乱码。
修复方案
有两种常用的修复方式,任选其一即可:
- 拿到响应后手动指定编码
在if request.status_code == 200:代码块下新增一行指定编码:
request.encoding = 'utf-8'
后续可直接用request.text传递给BeautifulSoup解析。
- 给BeautifulSoup显式指定解析编码
创建soup对象时增加编码参数:
soup = bs(request.content, features="html.parser", from_encoding='utf-8')
修改后的完整代码如下:
from bs4 import BeautifulSoup as bs import requests def main(): current_headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, " "like Gecko) Chrome/92.0.4515.159 Safari/537.36", 'Accept-Language': 'en-GB,jp,en;q=0.5', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,' 'image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'Accept-Encoding': 'gzip, deflate, br'} sess = requests.session() sess.headers.update(current_headers) resp = sess.get("https://nyaa.si/user/Mashin") if resp.status_code == 200: print("Successful in getting site") # 手动指定编码 resp.encoding = 'utf-8' print(resp.encoding) soup = bs(resp.text, features="html.parser") for table_row in soup.find_all('tr')[1:]: tds = table_row.find_all('td') print(tds[1].find("a", class_='')['title']) break if __name__ == '__main__': main()
运行后输出和预期结果一致:[마신] [2021.08.25] TVアニメ「ひぐらしのなく頃に 卒」EDテーマ「Missing Promise」/鈴木このみ [MP3 320K]
如果修改后依旧乱码,修改你运行代码的终端默认编码为UTF-8即可。
内容的提问来源于stack exchange,提问作者Moeed Azhar
相关产品推荐
相关产品推荐

