如何用Python抓取日文内容?解决抓取后内容乱码问题
网页爬取乱码修复方案
你的问题是编码不匹配导致的乱码,目标网站使用EUC-KR(韩文编码),但你强制用UTF-8解析,才出现了类似Á¶¼±°úÇбâ¼úÃѷøÍ的乱码字符。
Requests修复代码
直接使用原始字节流,指定正确编码解码:
headers = { 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'Accept-Language': 'en-US,en;q=0.9', 'Cache-Control': 'max-age=0', 'Connection': 'keep-alive', 'If-Modified-Since': 'Mon, 22 May 2017 16:51:07 GMT', 'If-None-Match': '"269-5501fad6c02b2-gzip"', 'Referer': 'http://kcna.co.jp/', 'Upgrade-Insecure-Requests': '1', 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/100.0.4896.127 Safari/537.36', } response = requests.get('http://kcna.co.jp/item2/2001/200107/news07/01.htm#10', headers=headers, verify=False) # 用EUC-KR解码原始响应字节 html_content = response.content.decode('euc-kr') resp = HtmlResponse(url='', body=html_content.encode('utf-8'), encoding='utf-8') print(resp.css('p::text').get())
Scrapy修复方案
在爬虫中强制指定响应编码:
def parse(self, response): # 手动设置响应编码为EUC-KR response.encoding = 'euc-kr' target_text = response.css('p::text').get() print(target_text)
核心要点
- 不要依赖requests/scrapy的自动编码猜测,该网站的实际编码是
EUC-KR,必须手动指定。 - 避免直接用
response.text(会自动猜测编码),改用response.content获取原始字节后解码,能避免编码猜测错误。
内容的提问来源于stack exchange,提问作者Umair Ayub
相关产品推荐
相关产品推荐

