从SQLite3批量获取URL爬取股票数据时出现Beautiful Soup NoneType错误
解决批量爬取股票报价时的BeautifulSoup NoneType错误
我明白你遇到的问题了——单条URL爬取股票价格完全正常,但从SQLite数据库批量读取URL循环处理时,却碰到了'NoneType' object has no attribute 'text'的报错。先理清楚问题出在哪,再一步步解决。
首先,你提供的单条爬取代码是可以正常运行的:
import requests from bs4 import BeautifulSoup as bs input_str = input("Insert stock's url:") if input_str=="": input_str="https://www.teleborsa.it/indici-italia/ftse-mib" res = requests.get(input_str) soup = bs(res.content,'lxml') price = soup.find("span", class_="h-price fc0").text print("Stock price ",input_str," è ",price)
但批量处理的代码里出现了错误,你给出的循环代码片段是:
# reading records for row in rows: input_str=row[6] res = request.get(input_str) soup = bs(res.content,'html.parser') price = soup.find("span", class_="h-price fc0").text curdata =soup.find("div", class_="header-bottom fc3").text
错误原因分析
这个报错的核心是:soup.find()没有找到你指定的元素,返回了None,而你直接调用了None.text,自然会触发AttributeError。具体可能的原因有几个:
- 低级拼写错误:你写的是
request.get(),少了个s,正确的应该是requests.get()——这个错误会直接导致请求失败,res.content是无效内容,BeautifulSoup解析后根本找不到目标元素。 - 页面结构不一致:数据库里的某些URL对应的页面,和你测试的单条页面结构不同,不存在
h-price fc0类的span或者header-bottom fc3类的div。 - 被网站反爬拦截:批量请求频率太高,网站返回了403/503这类错误页面,页面内容不是正常的股票详情页,自然找不到目标元素。
- URL无效:数据库里的某些URL格式错误、过期或者是空值。
解决方案
1. 先修复拼写错误
把request.get(input_str)改成requests.get(input_str),这是最容易忽略但影响最大的问题。
2. 增加完善的错误处理逻辑
在代码里加入检查和异常捕获,避免程序直接崩溃,还能定位具体哪个URL出了问题:
import requests from bs4 import BeautifulSoup as bs import time # 模拟浏览器请求头,避免被拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 假设你已经完成了数据库连接、查询,得到了rows for row in rows: input_str = row[6] # 先检查URL是否有效 if not input_str or not input_str.startswith(('http://', 'https://')): print(f"跳过无效URL: {input_str}") continue try: # 发送请求,设置超时避免卡住 res = requests.get(input_str, headers=headers, timeout=10) # 检查请求是否成功(比如返回404/500会抛出异常) res.raise_for_status() soup = bs(res.content, 'html.parser') # 先判断元素是否存在,再获取text price_elem = soup.find("span", class_="h-price fc0") price = price_elem.text.strip() if price_elem else "未找到价格数据" curdata_elem = soup.find("div", class_="header-bottom fc3") curdata = curdata_elem.text.strip() if curdata_elem else "未找到附加数据" print(f"股票 {input_str} 的价格: {price},附加信息: {curdata}") # 控制请求频率,避免被反爬 time.sleep(1) except Exception as e: print(f"处理URL {input_str} 时出错: {str(e)}")
3. 验证数据库中的URL
先把数据库里的所有URL打印出来,手动检查几个:
for row in rows: print(row[6])
看看有没有拼写错误、过期的链接,或者指向非目标页面的URL。
这样修改后,你的批量爬取代码应该就能稳定运行,而且能清晰定位每个URL的问题了。
内容的提问来源于stack exchange,提问作者abcoder
相关产品推荐
相关产品推荐

