You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests爬虫UTF-8编码下日语字符无法正常显示的问题求助

问题原因
  • requests默认对该站点的响应编码识别错误,识别结果为不支持东亚字符的ISO-8859-1,即使直接传递字节流给BeautifulSoup,库本身的自动编码检测也可能出现偏差,最终导致日语、韩文字符解析为乱码。
  • 如果你使用的Windows系统默认终端编码为GBK,也会在打印输出环节出现二次乱码。
修复方案

有两种常用的修复方式,任选其一即可:

  1. 拿到响应后手动指定编码
    在if request.status_code == 200:代码块下新增一行指定编码:
request.encoding = 'utf-8'

后续可直接用request.text传递给BeautifulSoup解析。

  1. 给BeautifulSoup显式指定解析编码
    创建soup对象时增加编码参数:
soup = bs(request.content, features="html.parser", from_encoding='utf-8')

修改后的完整代码如下:

from bs4 import BeautifulSoup as bs
import requests


def main():
    current_headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, "
                                     "like Gecko) Chrome/92.0.4515.159 Safari/537.36",
                       'Accept-Language': 'en-GB,jp,en;q=0.5',
                       'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,'
                                 'image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',
                       'Accept-Encoding': 'gzip, deflate, br'}
    sess = requests.session()
    sess.headers.update(current_headers)

    resp = sess.get("https://nyaa.si/user/Mashin")
    if resp.status_code == 200:
        print("Successful in getting site")
        # 手动指定编码
        resp.encoding = 'utf-8'
        print(resp.encoding)
        soup = bs(resp.text, features="html.parser")
        for table_row in soup.find_all('tr')[1:]:
            tds = table_row.find_all('td')
            print(tds[1].find("a", class_='')['title'])
            break


if __name__ == '__main__':
    main()

运行后输出和预期结果一致:[마신] [2021.08.25] TVアニメ「ひぐらしのなく頃に 卒」EDテーマ「Missing Promise」/鈴木このみ [MP3 320K]

如果修改后依旧乱码,修改你运行代码的终端默认编码为UTF-8即可。

内容的提问来源于stack exchange,提问作者Moeed Azhar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 04:09:04