爬取Yahoo Finance股票代码板块信息时遇AttributeError错误求助
解决Yahoo Finance股票板块信息爬取错误的方案
问题根源
出现AttributeError: 'NoneType' object has no attribute 'find_next'的核心原因是soup.find('span', text='Sector(s)')返回了None——要么请求被网站识别为爬虫,返回的页面内容不完整;要么页面HTML结构已更新,原选择器无法匹配目标元素。
修正方案
1. 添加请求头模拟浏览器访问
Yahoo Finance会拦截无标识的爬虫请求,添加User-Agent头可以让请求模拟浏览器行为,获取完整页面内容:
from bs4 import BeautifulSoup import requests url = 'https://finance.yahoo.com/quote/AAPL/profile?p=AAPL' # 模拟Chrome浏览器请求头,可根据自身浏览器版本调整 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } r = requests.get(url, headers=headers) soup = BeautifulSoup(r.content, 'html.parser') # 先确认能找到"Sector(s)"标签 sector_label = soup.find('span', string='Sector(s)') if sector_label: sector_element = sector_label.find_next('span', class_='Fw(600)') if sector_element: print(sector_element.text) # 预期输出:Technology else: print("无法定位板块内容元素") else: print("无法找到'Sector(s)'标签,页面结构可能已更新")
2. 适配更新后的页面结构(备选)
如果Yahoo Finance调整了页面HTML结构,原标签选择器失效,可以通过父容器定位板块信息:
from bs4 import BeautifulSoup import requests url = 'https://finance.yahoo.com/quote/AAPL/profile?p=AAPL' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } r = requests.get(url, headers=headers) soup = BeautifulSoup(r.content, 'html.parser') # 定位资产信息的父容器 profile_container = soup.find('div', class_='asset-profile-container') if profile_container: # 查找包含"Sector"的文本元素 sector_row = profile_container.find('div', string=lambda t: t and 'Sector' in t.strip()) if sector_row: # 获取相邻的内容元素 sector_content = sector_row.find_next_sibling('div').text.strip() print(sector_content) else: print("页面结构已更新,请检查元素选择器") else: print("无法找到资产信息容器")
注意事项
User-Agent可以从浏览器开发者工具中复制真实值,避免被网站拦截。- 网页结构可能随时间变化,若再次失效,需重新查看页面HTML,调整元素选择器。
内容的提问来源于stack exchange,提问作者Matiouz123
相关产品推荐
相关产品推荐

