雅虎财经股票Profile页行业信息爬虫报错及批量获取方案咨询
问题原因
- 未添加请求头伪装,雅虎财经默认拦截无合法UA标识的爬虫请求,返回的页面内容不包含浏览器端可见的完整DOM结构,导致
find_all('div',attrs={'id':'Main'})返回空列表,取索引[0]时触发越界错误。 data-reactid是React框架动态生成的属性值,会随页面版本、请求场景变化,用该属性定位元素稳定性极低。
修复后的单只股票抓取代码
import bs4 as bs import requests # 添加请求头伪装成浏览器 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } r = requests.get('https://finance.yahoo.com/quote/AAPL/profile?p=AAPL', headers=headers) soup = bs.BeautifulSoup(r.content, 'lxml') # 用固定标签文本定位,不受动态属性影响 industry_label = soup.find('span', string='Industry') if industry_label: industry = industry_label.find_next('span').get_text(strip=True) print(industry) else: print("未获取到行业信息")
批量抓取多只股票行业信息方案
- 原生爬虫实现:将需要查询的股票代码存入列表,循环拼接请求URL,每次请求后加1-2秒延时避免触发反爬,添加异常捕获逻辑跳过抓取失败的个股,示例逻辑如下:
import bs4 as bs import requests import time headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } tickers = ['AAPL', 'MSFT', 'GOOGL', 'AMZN', 'TSLA'] result = {} for ticker in tickers: try: url = f'https://finance.yahoo.com/quote/{ticker}/profile?p={ticker}' r = requests.get(url, headers=headers, timeout=10) soup = bs.BeautifulSoup(r.content, 'lxml') industry_label = soup.find('span', string='Industry') if industry_label: result[ticker] = industry_label.find_next('span').get_text(strip=True) else: result[ticker] = '无数据' # 控制请求频率 time.sleep(1.5) except Exception as e: result[ticker] = f'抓取失败:{str(e)}' print(result)
- 简化实现:可直接使用封装好雅虎财经接口的第三方库,无需自行处理页面解析和反爬基础逻辑,调用效率更高。
内容的提问来源于stack exchange,提问作者user16233760
相关产品推荐
相关产品推荐

