使用BeautifulSoup提取h2标签文本报NoneType不可迭代错误求助
问题修复方案
错误原因
- select方法使用错误:
soup.select()接收的是CSS选择器字符串,你传入class_参数属于find_all()的用法,当前写法无法匹配到你要的指定class的div,导致遍历到的大量div下不存在h2标签,x.find('h2')返回None,循环None就会抛出你遇到的类型错误。 - 冗余嵌套循环:你写了两层一模一样的遍历a标签的循环,属于完全不必要的重复逻辑,会重复发起大量无效请求。
- h2文本获取逻辑错误:
find('h2')返回的是单个Tag对象,不是可迭代的列表,不需要嵌套循环遍历,直接调用.text属性就能拿到文本。 - 缺少请求头:直接发起requests请求很容易被站点反爬拦截,返回异常页面导致无法匹配到元素。
修正后代码
import time import requests from bs4 import BeautifulSoup from csv import writer # 加请求头模拟浏览器访问,避免被反爬 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } base_url = "https://cognitiveclass.ai" r = requests.get(f"https://cognitiveclass.ai/search?q=data+science", headers=headers) soup = BeautifulSoup(r.text, 'html.parser') # 去掉冗余的第二层循环,直接遍历一次a标签即可 for a_tag in soup.find_all('a', href=True): print('-' * 60) href = a_tag['href'] print('href: ', href) # 避免请求外部链接,只爬取本站点路径 if not href.startswith('/'): continue request_href = requests.get(base_url + href, headers=headers) child_soup = BeautifulSoup(request_href.text, 'html.parser') # 正确匹配指定class的div div_title = child_soup.find_all('div', class_ = 'col-md-pull-4') for x in div_title: vendor = x.find('h2') # 先判断vendor是否存在,存在再取文本 if vendor: print(vendor.text.strip()) # 如果你只需要第一个匹配的结果可以保留break,不需要就删除 break
内容的提问来源于stack exchange,提问作者Abdul Ahad
相关产品推荐
相关产品推荐

