You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取雅虎Finance页面出现UnicodeDecodeError如何正确解码?

网页爬取内容解码错误排查及解决方案

错误原因

报错触发的核心原因是你请求的雅虎财经站点返回了gzip压缩格式的响应数据,0x8b是gzip压缩文件的特征字节,urllib默认不会自动解压压缩响应,你直接将压缩后的二进制数据用utf-8解码自然会触发编码错误。

可行解码方案

提供两种可直接使用的解决方案:

方案1:自动识别并解压压缩响应(兼容性最优)

先判断响应的压缩格式,解压后再做解码,适配所有开启了传输压缩的站点:

import urllib.request
import gzip
from bs4 import BeautifulSoup

site_url='https://finance.yahoo.com/quote/DUK?p=DUK'
r = urllib.request.urlopen(site_url)
raw_content = r.read()

# 识别压缩格式并处理
if r.getheader('Content-Encoding') == 'gzip':
    site_content = gzip.decompress(raw_content).decode('utf-8')
else:
    site_content = raw_content.decode('utf-8')

# 后续代码无需修改
with open('saved_page.html', 'w', encoding='utf-8') as f:
    f.write(site_content)

s = BeautifulSoup(site_content, 'html.parser')

注:写文件时建议显式指定encoding='utf-8',避免不同系统默认编码不一致导致的保存乱码问题

方案2:请求时声明不接受压缩(逻辑最简单)

发送请求时在请求头指定不接受压缩内容,让服务器直接返回明文HTML,省略解压步骤:

import urllib.request
from bs4 import BeautifulSoup

site_url='https://finance.yahoo.com/quote/DUK?p=DUK'
# 构造请求头,声明不接受压缩响应
req = urllib.request.Request(
    site_url,
    headers={'Accept-Encoding': 'identity'}
)
r = urllib.request.urlopen(req)
site_content = r.read().decode('utf-8')

# 后续代码无需修改
with open('saved_page.html', 'w', encoding='utf-8') as f:
    f.write(site_content)

s = BeautifulSoup(site_content, 'html.parser')

内容的提问来源于stack exchange,提问作者Anibal Ceballos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 12:54:03