Python Requests爬雅虎财经数据偶发成功多数返回null如何解决
雅虎财经流通股爬取timeSeries为null问题修复
这个问题很多爬取雅虎财经的开发者都遇到过,QUOTE_TIMESERIES_LOAD错误本质是请求没有通过雅虎的校验,时序数据没有正常加载,不是单独的解析逻辑问题。
核心触发原因
- 请求头缺失关键字段,仅携带User-Agent会被雅虎的反爬规则识别为爬虫请求,拒绝返回时序数据
- 没有维持会话Cookie,首次访问雅虎站点会下发校验Cookie,不带Cookie直接请求个股页面大概率拿不到动态加载的时序数据
- 请求频率过高触发限流,雅虎对单IP的短时间请求次数有明确阈值,触发后会直接拦截动态数据返回
- 原有字符串截断的解析逻辑鲁棒性差,不同地区、不同个股的页面JS结构有细微差异,硬编码split规则容易截错JSON内容
- 部分个股(如粉单、退市标的)本身不公开对应时间粒度的流通股数据,这种属于正常返回空值
可落地修复方案
- 补全必带请求头,不要只传User-Agent,参考浏览器真实请求补全Accept、Accept-Language、Referer字段
- 用
requests.Session()维持会话,先访问雅虎财经首页拿到合法Cookie后再请求个股页面,不要裸发请求 - 控制请求间隔,单IP两次请求间隔不低于2秒,触发限流后暂停10-15分钟再重试,不要高频撞墙
- 替换硬编码的字符串截断逻辑,用正则精准匹配
root.App.main挂载的JSON对象,避免页面结构微调导致解析失败 - 增加错误重试逻辑,第一次返回加载错误时等待3-5秒重试一次,能解决偶发的网络波动、节点调度导致的加载失败
修复后参考代码
import requests import json import re import time def my_parse_json_outstanding_shares(stock_code): # 用Session维持Cookie session = requests.Session() headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/15.5 Safari/605.1.15', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', 'Accept-Language': 'zh-CN,zh-Hans;q=0.9,en;q=0.8', 'Referer': 'https://finance.yahoo.com/', 'Connection': 'keep-alive' } # 先访问首页拿有效Cookie session.get('https://finance.yahoo.com', headers=headers, timeout=10) time.sleep(1.5) target_url = f'https://finance.yahoo.com/quote/{stock_code}/shares-outstanding' resp = session.get(target_url, headers=headers, timeout=10) resp.raise_for_status() # 正则精准匹配挂载的JSON数据,替代硬编码split json_match = re.search(r'root\.App\.main\s*=\s*(\{.*?\});\s*\(this\)', resp.text, re.DOTALL) if not json_match: return {} try: app_data = json.loads(json_match.group(1)) ts_store = app_data['context']['dispatcher']['stores']['QuoteTimeSeriesStore'] # 遇到加载错误重试一次 if ts_store.get('errorList') and 'QUOTE_TIMESERIES_LOAD' in ts_store['errorList']: time.sleep(4) resp = session.get(target_url, headers=headers, timeout=10) json_match = re.search(r'root\.App\.main\s*=\s*(\{.*?\});\s*\(this\)', resp.text, re.DOTALL) if not json_match: return {} app_data = json.loads(json_match.group(1)) ts_store = app_data['context']['dispatcher']['stores']['QuoteTimeSeriesStore'] # 保留原有raw字段清洗逻辑 cleaned_data = json.dumps(ts_store).replace('{}', 'null') cleaned_data = re.sub(r'\{[\'"]raw[\'"]:(.*?),(.*?)\}', r'\1', cleaned_data) return json.loads(cleaned_data) except Exception: return {}
注意:如果连续3次以上重试都返回
QUOTE_TIMESERIES_LOAD错误,基本可以判定是IP被临时限流,换代理IP或者等待15分钟后再请求即可,不要持续高频请求,不然会拉长封禁时间。
内容的提问来源于stack exchange,提问作者Nicolas
相关产品推荐
相关产品推荐

