You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python抓取日文内容?解决抓取后内容乱码问题

网页爬取乱码修复方案

你的问题是编码不匹配导致的乱码,目标网站使用EUC-KR(韩文编码),但你强制用UTF-8解析,才出现了类似Á¶¼±°úÇбâ¼úÃѷøÍ的乱码字符。

Requests修复代码

直接使用原始字节流,指定正确编码解码:

headers = {
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',
    'Accept-Language': 'en-US,en;q=0.9',
    'Cache-Control': 'max-age=0',
    'Connection': 'keep-alive',
    'If-Modified-Since': 'Mon, 22 May 2017 16:51:07 GMT',
    'If-None-Match': '"269-5501fad6c02b2-gzip"',
    'Referer': 'http://kcna.co.jp/',
    'Upgrade-Insecure-Requests': '1',
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/100.0.4896.127 Safari/537.36',
}

response = requests.get('http://kcna.co.jp/item2/2001/200107/news07/01.htm#10', headers=headers, verify=False)
# 用EUC-KR解码原始响应字节
html_content = response.content.decode('euc-kr')
resp = HtmlResponse(url='', body=html_content.encode('utf-8'), encoding='utf-8')
print(resp.css('p::text').get())

Scrapy修复方案

在爬虫中强制指定响应编码:

def parse(self, response):
    # 手动设置响应编码为EUC-KR
    response.encoding = 'euc-kr'
    target_text = response.css('p::text').get()
    print(target_text)

核心要点

  • 不要依赖requests/scrapy的自动编码猜测,该网站的实际编码是EUC-KR,必须手动指定。
  • 避免直接用response.text(会自动猜测编码),改用response.content获取原始字节后解码,能避免编码猜测错误。

内容的提问来源于stack exchange,提问作者Umair Ayub

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 13:42:48