Python使用BeautifulSoup多页爬虫迭代时出现NoneType不可迭代问题
问题原因分析
- 原因1:
find()方法匹配不到对应元素时会返回None,直接对None调用属性、执行遍历操作就会抛出'NoneType' object is not iterable报错。你测试的页面均存在class="ad-price"的元素,但部分页面不存在你写的另外两类元素,所以只有价格字段能正常执行。 - 原因2:提取标题时你错写了属性名,
stripped_string是单数属性,返回的是单个字符串(不存在则返回None),不能直接用于for循环遍历,正确的多文本提取属性是复数的stripped_strings。 - 原因3:提取地址时你直接遍历
find()返回的Tag对象,会迭代它的所有子节点,而非你需要的地址文本内容,逻辑不符合需求。 - 额外隐患:如果某页某个字段缺失,你直接跳过不追加值,会导致三个列表最终长度不一致,无法正常拼接为DataFrame。
修复代码
import pandas as pd from requests import get from bs4 import BeautifulSoup precio = [] direccion = [] titulo = [] for url in urls.values: url_ = url[0] response = get(url_) page = BeautifulSoup(response.content, 'html.parser') # 提取价格 price_tag = page.find("span", attrs={"class" : "ad-price"}) price_text = ''.join(price_tag.stripped_strings) if price_tag else '' precio.append(price_text) # 提取标题 title_tag = page.find("div", attrs={"class" : "title"}) title_text = ''.join(title_tag.stripped_strings) if title_tag else '' titulo.append(title_text) # 提取地址 loc_tag = page.find("div", attrs = {"class" : "location-name"}) loc_text = ''.join(loc_tag.stripped_strings) if loc_tag else '' direccion.append(loc_text) # 生成最终DataFrame df = pd.DataFrame({ 'precio': precio, 'titulo': titulo, 'direccion': direccion })
如果修复后仍有字段为空,可打开对应页面的开发者工具检查元素,确认你写的class名称是否和页面实际的class属性完全一致,部分网站的class可能存在拼写、大小写或者动态拼接的差异。
内容的提问来源于stack exchange,提问作者Ramiro Guzmán
相关产品推荐
相关产品推荐

