You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取Morningstar金融数据时第12个标的固定返回页面丢失问题

爬取Morningstar金融数据固定在第12个代码报错问题排查

问题描述

使用股票代码列表爬取Morningstar的金融数据,运行到第12个代码时固定返回页面丢失。对应页面实际存在,即便将可正常访问的代码换到第12位,仍会出现相同报错,已尝试添加时间延迟、循环重试出错代码的方案,均无法解决。

相关代码

核心测试代码

for symbol in symbols:   
    url_morningstar = 'https://www.morningstar.com/funds/xnas/{}/quote'
    response = requests.get(url_morningstar.format(symbol))
    mySoup = BeautifulSoup(response.text, 'html.parser')
    htmlData = mySoup.findAll('span',{'class':'mdc-data-point mdc-data-point--number'})
    while(len(htmlData) == 0):
            print(symbol, ' ---ERROR---')
            print(htmlData)
            #print(response.text)
            response = requests.get(url_morningstar.format(symbol))
            mySoup = BeautifulSoup(response.text, 'html.parser')
            htmlData = mySoup.findAll('span',{'class':'mdc-data-point mdc-data-point--number'})
    duration = htmlData[-1].text.strip()
    nav = htmlData[0].text.strip()

完整可复现代码

symbols = []
with open('symbols.csv') as csvfile:
    reader = csv.reader(csvfile)
    for row in reader:
        symbols.append(row[0])
full_data = []
for symbol in symbols:   
    print(symbol) 
    url_schwab = 'https://www.schwab.wallst.com/Prospect/Research/MutualFunds/Summary.asp?symbol={}'
    url_morningstar = 'https://www.morningstar.com/funds/xnas/{}/quote'

    response = requests.get(url_schwab.format(symbol))
    mySoup = BeautifulSoup(response.text, 'html.parser')
    table = mySoup.find('div',{'id':'detailsWrapper'})
    rows = table.findAll('table',{'class':'tableType1'})
    headers = []
    output = []
    schwab_dict = {}
    for row in rows:
        cols = row.find('tbody').find('tr').findAll('td')
        colNames = row.find('tbody').find('tr').findAll('th')
        colNames = [ele.text.strip() for ele in colNames]
        cols = [ele.text.strip() for ele in cols]

        output.append([ele for ele in cols if ele]) 
        headers.append([ele for ele in colNames if ele])
    headers[1] = ['YTD Return']
    headers[6] = ['Distribution Yield']
    for i in range(len(headers)):
        schwab_dict[headers[i][0]] = output[i][0]


    response = requests.get(url_morningstar.format(symbol))
    mySoup = BeautifulSoup(response.text, 'html.parser')
    htmlData = mySoup.findAll('span',{'class':'mdc-data-point mdc-data-point--number'})
    while(len(htmlData) == 0):
        print(symbol, ' ---ERROR---')
        print(htmlData)
        #print(response.text)
        response = requests.get(url_morningstar.format(symbol))
        mySoup = BeautifulSoup(response.text, 'html.parser')
        htmlData = mySoup.findAll('span',{'class':'mdc-data-point mdc-data-point--number'})
    duration = htmlData[-1].text.strip()
    nav = htmlData[0].text.strip()

    # extract Duration EXP ratio YTD 2021 SEC Yield Price Last Updated
    results = [duration, schwab_dict['Net Expense Ratio'], schwab_dict['YTD Return'], schwab_dict['30-Day SEC Yield'], nav ]
    full_data.append([results])
with open('scrappedData.csv', 'x') as csvfile:
    writer = csv.reader(csvfile)
    writer.writerow(full_data)

symbols.csv内容

DLSNX
FFRHX
MWLDX
OSTIX
PRWBX
VBIRX
VSGBX
VFSTX
VFISX
FSTFX
PRFSX
VMLTX 
VWSTX
DODIX
DLTNX
FAGIX
SPHIX
FTHRX
FBNDX
FADMX
FTBFX
LSBRX
MWTRX
RPSIX
VFIIX
VWEHX
VBILX
VFICX
VFITX
VBTLX
FLTMX
PRSMX
VCAIX
VWITX
PRPIX
PRULX
VIPSX
VBLAX
VWESX
VUSTX
FCTFX
FHIGX
FTFMX
FTABX
PRINX
PRFHX
PRTAX
VCITX
VWAHX
VWLTX
LSGLX
RPIBX
VTABX
FCVSX
VWINX

问题原因

首先观察第12个代码VMLTX后面带有多余空格,请求时会被拼接为https://www.morningstar.com/funds/xnas/VMLTX /quote,空格会被编码为%20,直接导致请求路径错误,这是最直接的触发原因。

如果去除空格后仍出现固定位置报错,就是反爬机制生效:

  1. 默认requests发起的请求UA标识会被Morningstar直接识别为爬虫,短时间多次请求后就会触发频率限制,返回403或者空页面,这也是为什么更换代码到第12位仍然报错的核心原因:反爬基于请求频率和请求特征识别,和具体请求的股票代码无关。
  2. 没有使用会话保持,每次请求都是独立的,没有携带之前请求返回的cookie,也会触发反爬策略拦截。
  3. 现有重试逻辑是无延迟死循环重试,一旦触发反爬,会持续发送无效请求,反而会拉长限制时间。

修复建议

  • 处理symbols列表时给每个代码做strip()处理,去掉前后空白字符:symbols.append(row[0].strip())
  • 给所有请求添加模拟浏览器的UA头,示例:
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
response = requests.get(url, headers=headers)
  • 使用requests.Session()发起请求,自动保持cookie,提升请求合法性。
  • 每次请求之间添加1-3秒的随机延迟,模拟真人操作间隔。
  • 重试逻辑添加最大重试次数,避免死循环,重试前等待更长时间。

内容的提问来源于stack exchange,提问作者Ben Cole

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 19:45:04