使用Beautiful Soup获取Stocktwits数据时遇AttributeError问题求助
解决Stocktwits帖子数量爬取的AttributeError错误
问题场景
编写了一段用Beautiful Soup获取Stocktwits指定股票(如TSLA)帖子数量的Python代码,运行时触发AttributeError,提示'NoneType' object has no attribute 'text',错误出现在获取soup.find找到的span元素的text属性时——因为该方法未找到目标元素,返回了None。
代码如下:
def get_stocktwits_posts(symbol): url = f"https://stocktwits.com/symbol/{symbol}" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") post_count = soup.find("span", class_="st_1Zg91_0").text return int(post_count.replace(",", "")) symbol='TSLA' current_posts = get_stocktwits_posts(symbol)
错误详情:
AttributeError: 'NoneType' object has no attribute 'text'
错误原因
- 目标元素CSS类名已更新:网站可能调整了前端样式,
st_1Zg91_0这个类名不再对应帖子数量的span元素。 - 页面动态加载:帖子数量可能通过JavaScript动态渲染,直接用requests获取的静态HTML中没有该元素。
- 反爬拦截:requests请求未携带必要请求头,被网站识别为爬虫,返回的内容不是正常页面。
解决方法
1. 先判断元素是否存在,避免直接访问text
修改代码,先检查soup.find的返回值,防止None调用text属性:
def get_stocktwits_posts(symbol): url = f"https://stocktwits.com/symbol/{symbol}" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") post_elem = soup.find("span", class_="st_1Zg91_0") if not post_elem: # 元素未找到时的处理逻辑,比如返回0或抛出明确异常 return 0 post_count = post_elem.text return int(post_count.replace(",", "")) symbol='TSLA' current_posts = get_stocktwits_posts(symbol)
2. 重新获取正确的CSS选择器
打开浏览器访问Stocktwits股票页面,按F12打开开发者工具,找到帖子数量对应的元素,复制最新的CSS类名或其他定位方式(比如通过父元素层级定位),替换代码中的class参数即可。
3. 处理动态加载和反爬
- 添加请求头:模拟浏览器请求,避免被拦截:
def get_stocktwits_posts(symbol): url = f"https://stocktwits.com/symbol/{symbol}" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 后续逻辑同上
- 使用动态渲染工具:如果页面是JS动态加载的,改用selenium获取渲染后的页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager def get_stocktwits_posts(symbol): url = f"https://stocktwits.com/symbol/{symbol}" driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get(url) post_elem = driver.find_element(By.CLASS_NAME, "st_1Zg91_0") post_count = post_elem.text driver.quit() return int(post_count.replace(",", ""))
内容的提问来源于stack exchange,提问作者abcoder
相关产品推荐
相关产品推荐

