You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从SQLite3批量获取URL爬取股票数据时出现Beautiful Soup NoneType错误

解决批量爬取股票报价时的BeautifulSoup NoneType错误

我明白你遇到的问题了——单条URL爬取股票价格完全正常,但从SQLite数据库批量读取URL循环处理时,却碰到了'NoneType' object has no attribute 'text'的报错。先理清楚问题出在哪,再一步步解决。

首先,你提供的单条爬取代码是可以正常运行的:

import requests
from bs4 import BeautifulSoup as bs
input_str = input("Insert stock's url:")
if input_str=="":
    input_str="https://www.teleborsa.it/indici-italia/ftse-mib"
res = requests.get(input_str)
soup = bs(res.content,'lxml')
price = soup.find("span", class_="h-price fc0").text
print("Stock price ",input_str," è ",price)

但批量处理的代码里出现了错误,你给出的循环代码片段是:

# reading records
for row in rows:
    input_str=row[6]
    res = request.get(input_str)
    soup = bs(res.content,'html.parser')
    price = soup.find("span", class_="h-price fc0").text
    curdata =soup.find("div", class_="header-bottom fc3").text

错误原因分析

这个报错的核心是:soup.find()没有找到你指定的元素,返回了None,而你直接调用了None.text,自然会触发AttributeError。具体可能的原因有几个:

  • 低级拼写错误:你写的是request.get(),少了个s,正确的应该是requests.get()——这个错误会直接导致请求失败,res.content是无效内容,BeautifulSoup解析后根本找不到目标元素。
  • 页面结构不一致:数据库里的某些URL对应的页面,和你测试的单条页面结构不同,不存在h-price fc0类的span或者header-bottom fc3类的div。
  • 被网站反爬拦截:批量请求频率太高,网站返回了403/503这类错误页面,页面内容不是正常的股票详情页,自然找不到目标元素。
  • URL无效:数据库里的某些URL格式错误、过期或者是空值。

解决方案

1. 先修复拼写错误

把request.get(input_str)改成requests.get(input_str),这是最容易忽略但影响最大的问题。

2. 增加完善的错误处理逻辑

在代码里加入检查和异常捕获,避免程序直接崩溃,还能定位具体哪个URL出了问题:

import requests
from bs4 import BeautifulSoup as bs
import time

# 模拟浏览器请求头,避免被拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

# 假设你已经完成了数据库连接、查询,得到了rows
for row in rows:
    input_str = row[6]
    # 先检查URL是否有效
    if not input_str or not input_str.startswith(('http://', 'https://')):
        print(f"跳过无效URL: {input_str}")
        continue
    
    try:
        # 发送请求,设置超时避免卡住
        res = requests.get(input_str, headers=headers, timeout=10)
        # 检查请求是否成功(比如返回404/500会抛出异常)
        res.raise_for_status()
        
        soup = bs(res.content, 'html.parser')
        
        # 先判断元素是否存在,再获取text
        price_elem = soup.find("span", class_="h-price fc0")
        price = price_elem.text.strip() if price_elem else "未找到价格数据"
        
        curdata_elem = soup.find("div", class_="header-bottom fc3")
        curdata = curdata_elem.text.strip() if curdata_elem else "未找到附加数据"
        
        print(f"股票 {input_str} 的价格: {price},附加信息: {curdata}")
        # 控制请求频率,避免被反爬
        time.sleep(1)
    
    except Exception as e:
        print(f"处理URL {input_str} 时出错: {str(e)}")

3. 验证数据库中的URL

先把数据库里的所有URL打印出来,手动检查几个:

for row in rows:
    print(row[6])

看看有没有拼写错误、过期的链接,或者指向非目标页面的URL。

这样修改后,你的批量爬取代码应该就能稳定运行,而且能清晰定位每个URL的问题了。

内容的提问来源于stack exchange,提问作者abcoder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 19:52:44