如何解决BeautifulSoup解析重复Class的问题
解决方法
针对class重复导致匹配到空文本元素的问题,你可以通过以下几种方式精准定位到目标价格元素:
方法1:遍历所有匹配元素,筛选非空文本
用find_all获取所有对应class的span元素,然后筛选出包含有效价格文本(带£符号)的元素:
url = 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/tiffany-t-t1-circle-pendant-69901190/' req = Request( url=url, headers={'User-Agent': 'Mozilla/5.0'} ) webpage = urlopen(req, context=ctx).read() soup = bs4.BeautifulSoup(webpage, "html.parser") # 获取所有匹配class的span元素 price_elements = soup.find_all("span", class_="product-description__addtobag_btn_text-static_price-wrapper_price") # 遍历筛选有价格文本的元素 for elem in price_elements: price_text = elem.get_text(strip=True) if price_text and "£" in price_text: print(price_text) break # 找到目标后停止遍历
方法2:定位到价格所在的父容器,缩小匹配范围
如果目标价格在特定的父元素内,可以先定位父容器,再在其中查找价格元素,避免匹配到无关的同class元素:
# 先定位价格所在的父容器(根据实际HTML结构调整class) parent_container = soup.find("div", class_="product-description__addtobag_btn_text-static") if parent_container: target_price = parent_container.find("span", class_="product-description__addtobag_btn_text-static_price-wrapper_price") if target_price: print(target_price.get_text(strip=True))
方法3:使用find_next获取下一个同class元素
如果第一个匹配元素是空的,目标元素是它的下一个同class兄弟元素,可以用find_next直接定位:
first_empty_price = soup.find("span", class_="product-description__addtobag_btn_text-static_price-wrapper_price") if first_empty_price: target_price = first_empty_price.find_next("span", class_="product-description__addtobag_btn_text-static_price-wrapper_price") if target_price: print(target_price.get_text(strip=True))
这些方法都能避开第一个空文本的元素,精准拿到£920的价格。
内容的提问来源于stack exchange,提问作者Seedizens
相关产品推荐
相关产品推荐

