爬取供应商新品页面时遇'NoneType' object has no attribute 'text'错误求助
解决爬取供应商新品时的'NoneType' object has no attribute 'text'错误
嘿,这个问题我太熟悉了!报错的根源很简单:你写代码的时候默认了所有带有po_blok类的div容器里,一定存在po_blok_stoc这个span元素,但实际网页里肯定有部分产品条目不符合这个假设——比如有些产品可能没显示库存状态,或者页面结构和你拿到的目标片段不一样。当result.find(...)找不到对应元素时,会返回None,这时候再去调用.text,自然就触发了AttributeError。
两种靠谱的解决办法:
核心思路都是先确认元素存在,再去获取文本内容,同时给缺失的情况设置默认值。
方法1:先存元素再判断(最清晰)
把查找元素的结果先存起来,再判断是否为None,这样还能避免重复调用find方法:
# import libraries from urllib.request import urlopen as uReq from bs4 import BeautifulSoup # specify the url url = "https://www.erotischegroothandel.nl/nieuweproducten/" # Connect to the website and return the html to the variable ‘page’ uClient = uReq(url) page_html = uClient.read() uClient.close() # parse the html using beautiful soup and store in variable `soup` soup = BeautifulSoup(page_html, 'html.parser') results = soup.find_all('div', {'class': 'po_blok'}) records = [] for result in results: titel = result.find('span', {'class': 'po_blok_titl'}).text staat = result.find('span', {'class': 'po_blok_nieu'}).text # 先获取库存元素,再判断是否存在 voorraad_element = result.find('span', {'class': 'po_blok_stoc'}) # 存在就取文本,不存在就用默认值 voorraad = voorraad_element.text.strip() if voorraad_element else "无库存信息" records.append((titel, staat, voorraad)) print(records)
方法2:用一行代码的三元表达式(更简洁)
如果觉得单独存元素麻烦,也可以用一行搞定,不过可读性稍弱一点:
# 替换原代码中的voorraad行 voorraad = result.find('span', {'class': 'po_blok_stoc'}).text.strip() if result.find('span', {'class': 'po_blok_stoc'}) else "无库存信息"
额外排查小技巧:
如果你想搞清楚到底是哪些产品没有库存元素,可以在循环里加个打印,看看每个条目的结构:
for idx, result in enumerate(results): print(f"\n=== 第{idx+1}个产品的结构 ===") print(result.prettify())
这样就能直观看到哪些po_blok里没有po_blok_stoc,方便你后续优化爬取逻辑。
内容的提问来源于stack exchange,提问作者Tim Bronsgeest
相关产品推荐
相关产品推荐

