使用requests.get获取Yahoo Finance页面未输出HTML,求排查
问题原因及解决办法
核心原因
Yahoo Finance返回的响应内容采用了gzip压缩,尽管requests默认会自动处理压缩响应,但部分场景下可能因环境或响应头细节问题,导致自动解压缩失效,直接打印response.text就会出现乱码。
解决步骤
方法1:明确指定压缩格式,确保自动解压缩
在请求头中添加Accept-Encoding字段,告知服务器接受gzip压缩格式,同时显式确认requests的自动解压缩逻辑:
import requests from bs4 import BeautifulSoup url = "https://finance.yahoo.com/quote/AMZN" headers = { 'USER-AGENT': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", 'Accept-Encoding': 'gzip, deflate' } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'lxml') print(soup.prettify())
方法2:手动解码压缩内容
如果自动解压缩失效,可直接获取二进制响应内容,手动判断并解码:
import requests import gzip from bs4 import BeautifulSoup url = "https://finance.yahoo.com/quote/AMZN" headers = { 'USER-AGENT': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) # 检查响应是否为gzip压缩格式,手动解码 if 'gzip' in response.headers.get('Content-Encoding', ''): decoded_content = gzip.decompress(response.content).decode('utf-8') else: decoded_content = response.text soup = BeautifulSoup(decoded_content, 'lxml') print(soup.prettify())
额外注意事项
- 确保
requests版本为最新,旧版本可能存在压缩处理bug,执行pip install --upgrade requests即可更新。 - Yahoo Finance有反爬机制,频繁请求会被限制,建议添加请求间隔(如
time.sleep(2)),避免触发反爬策略。
内容的提问来源于stack exchange,提问作者Dimbo123
相关产品推荐
相关产品推荐

