使用BeautifulSoup爬虫时Python循环报错:'NoneType' object is not subscriptable
解决新蛋特惠页面爬取循环中的NoneType错误
嘿,我来帮你搞定这个爬取问题!你遇到的'NoneType' object is not subscriptable错误,本质是循环里的某些商品结构和你预期的不一样,导致代码访问了不存在的元素。
错误原因分析
具体到你的代码,问题出在brand= titleco[0].img["title"]这一行:
- 要么是某个商品的
item-container里没有item-branding这个div,导致titleco是空列表,取titleco[0]直接报错; - 要么是
item-branding存在,但里面没有<img>标签(比如品牌是用纯文本显示的),这时titleco[0].img返回None,再访问["title"]就会触发NoneType错误。
而且你的价格部分也没有做异常处理,万一某个商品没有price-current这个元素,同样会报错中断循环。
修改后的代码(带容错处理)
下面是加了完整判断的代码,能应对结构不一致的商品,保证循环正常执行:
containers = pagesoup.findAll("div",{"class":"item-container"}) for con in containers: # 获取商品标题,处理无img标签的情况 title = con.img["title"] if con.img else "无标题" # 获取品牌,兼容图片和文本两种展示方式 titleco = con.findAll("div",{"class":"item-branding"}) brand = "无品牌" if titleco: brand_element = titleco[0].img if brand_element: brand = brand_element["title"] else: # 提取纯文本品牌(有的商品品牌是文字不是图片) brand_text = titleco[0].get_text(strip=True) brand = brand_text if brand_text else "无品牌" # 获取价格,处理无价格元素的情况 priceco = con.findAll("li",{"class":"price-current"}) price = "无价格" if priceco: price = priceco[0].text.strip() # 打印结果,也可以存到列表/文件里 print(f"标题: {title}\n品牌: {brand}\n价格: {price}\n---")
另一种方案:用try-except捕获异常
如果你更倾向于简洁的代码,也可以用异常捕获来跳过有问题的商品,同时保留错误信息方便调试:
containers = pagesoup.findAll("div",{"class":"item-container"}) for idx, con in enumerate(containers): try: title = con.img["title"] titleco = con.findAll("div",{"class":"item-branding"}) brand = titleco[0].img["title"] priceco = con.findAll("li",{"class":"price-current"}) price = priceco[0].text.strip() print(f"商品 {idx+1}:\n标题: {title}\n品牌: {brand}\n价格: {price}\n---") except Exception as e: print(f"商品 {idx+1} 处理出错: {str(e)}") # 可选:打印出错商品的HTML结构,方便排查问题 # print(con.prettify()) continue
额外建议
爬取电商页面时,网站的HTML结构经常会微调,建议你:
- 用浏览器开发者工具(F12)检查几个有问题的商品,看看它们的品牌/价格是怎么展示的;
- 尽量避免直接用索引(比如
[0])访问元素,先判断列表是否为空; - 对于可能变化的元素,多考虑几种展示情况(比如品牌是图片还是文本)。
内容的提问来源于stack exchange,提问作者hamza mhadhbi
相关产品推荐
相关产品推荐

